Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claim Rejections - 35 USC § 112
The]following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 3-5, 8-22 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 1, 16, 19 recite the limitation "a modified transcription…wherein the modified transcription is a corrected transcription.” Are the modified and corrected transcription distinct, or does a modified transcript generate a corrected transcription, or vice versa. For the purposes of the art rejection infra Examiner will presume that the modified and corrected transcriptions are the same. The possible redundancy renders the claim indefinite and further confuses the scope of and renders indefinite dependent claims 3-5, 12, 15-18 which recite “the corrected transcription,” as well as claim 13 which recites “the modified transcription.” The remaining claims do not remedy and are similarly indefinite. Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 3-5, 7-22 rejected under 35 U.S.C. 103 as being unpatentable over Zhu: 20210375289 (of record) further in view of Ma: “CAN GENERATIVE LARGE LANGUAGE MODELS PERFORM ASR ERROR CORRECTION?” (copy provided by Examiner; copyright 7/9/2023; and hereinafter Ma).
Regarding claim 1
Zhu teaches:
A user electronic device comprising:
one or more microphones configured to capture raw audio data (Zhu: ¶ 541-543; Fig 9: plural systems comprising one or more microphones for capturing audio data such as for transcription processing, generation of text thereby, etc.); and
one or more processors and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations (Zhu: ¶ 40; Fig 1: processors operative of memory borne instructions such as to operate a meeting assistant functionality in concert with microphone(s)) comprising:
receiving the raw audio data captured by the one or more microphones (Zhu: ¶ 541-543; Fig 5, 9: microphone(s) record speech of participants for use downstream such as in the form of generated text, etc.);
processing the raw audio data using a speech transcriber to generate a live transcription of the raw audio data that comprises a plurality of text tokens (Zhu: ¶ 72, 82-84, 454, 541-544; Fig 9: such as by generation of speech to text, tagging, labelling, tokenization, etc. in concert with a language model, and other modalities, such as to make the recorded audio pliant to downstream processing such as by summarization, etc.);
processing the raw audio data to generate a speaker identification output that identifies, for each of the text tokens, a respective speaker for each of the text tokens in the live transcription (Zhu: ¶ 72, 79-84, 443, 541-543; Fig 1, 9: extracted audio features, such as a voice profile used to determine identity of a speaker);
generating an input text by modifying the live transcription to insert text identifying the respective speakers for each of the text tokens in the live transcription (Zhu: ¶ 36, 55, 72, 79-84, 443, 541-543; Fig 5, 9: transcription modified to include speaker identification); and
processing the input text generated from the live transcription using a language model neural network to generate a modified transcription (Zhu: ¶ 36, 55, 72, 79-84, 459-465, 468-477, 541-543; Fig 3-5: transcription processed for summarization such as by a transformer model) such as based on an instruction to correct the live transcription and wherein the modified transcription is a corrected transcription that corrects transcription errors in the live transcription (Zhu: ¶ 57, 68, 454, 490-492, 548, etc.; Fig 3: post processing includes instructions to correct errors within the transcription, such as grammatical errors, errors introduced by automatic speech recognition, etc.; such as to make a transcript more readable, more correct, etc.).
Zhu does not explicitly teach the correction performed by a language model neural network; that the first input comprises a single input sequence comprising a combination of both (i) a first input prompt that comprises an instruction to correct the live transcription and (ii) the input text (Ma: § 3.1; Fig 2: ; nor that the model generates the modified transcription as a corrected version of the input text, which corrects transcription errors by applying the instruction from the first input prompt to the input text.
In a related field of endeavor Ma teaches a system and method for using an LLM to perform error correction of automatic speech recognition results (Ma: Abstract; § 3: ASR error correction using an LLM); wherein a first input comprises a single input sequence comprising a combination of both (i) a first input prompt that comprises an instruction to correct the live transcription and (ii) the input text (Ma: ¶ 3.1; Fig 2: system asks or prompts “ChatGPT to directly output the corrected hypothesis without adding further explanation”; such as depicted variously in figure 2); wherein the model generates the modified transcription as a corrected version of the input text (Ma: § 4.2; Fig 2: “ChatGPT is effective at detecting errors in the given ASR hypotheses and generating the corrected transcription,” such as depicted in the figure).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to utilize the Ma taught LLM to improve the performance of transcription and automatic speech recognition by the Zhu taught electronic device by correction in the manner claimed and for at least the purpose of creating a training free system by which a device corrects a labelled transcript such as that of Zhu using a single correction prompt such as that taught by Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 3
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 1, the operations further comprising: processing a second input comprising a second prompt for a text analysis task and context data comprising the corrected transcription and using the language model neural network to generate a text output for the text analysis task for the corrected transcription (Zhu: ¶ 65: such as for providing similar functionality for additional meetings, related parties or instances of speech recognition); (Ma: Fig 2: system utilizes additional prompts to correct, edit, curate, etc. the summary of additional transcriptions). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 4
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 3, wherein the second prompt comprises an instruction to identify action items for a particular speaker, wherein the action items comprise (1) questions for the speaker to answer, (2) tasks for the speaker to complete, or both, and the text output comprises text derived from the corrected transcription that identifies one or more action items for the particular speaker (Zhu: ¶ 65, 72, 82-84, 454, 489, 541-544; Fig 5, 9: speakers, roles, etc. identified, tasks ascribed thereto; in the case of the figure 5 illustration the transcriptions apply to product design and the summary one for the industrial designer speaker, role, etc.). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 5
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 3, wherein the second prompt comprises an instruction to summarize the corrected transcription and the text output comprises text that summarizes the corrected transcription (Zhu: ¶ 462: such as by generating a meeting summary, etc.); (Ma: Abstract: such as additionally using the taught LLM for an additional prompt). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 7
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 3, wherein the language model neural network provides a summary of the audio data at predetermined time intervals (Zhu: ¶ 77, 207, 448, 462, 523-525: such as during a particular meeting by providing streaming speech, real time correction and summarization). Examiner takes official notice that periodic or scheduled summary of an ongoing or real time transcription was well known in the art before the effective filing date of the instant invention and would have comprised an obvious inclusion for at least the purpose of determining particular times for particular topics or relevance criteria such as a time duration of a speech, meeting, etc. The claim is thus considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 8
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 7, the operations further comprising: determining that a specified time interval has elapsed since a prior live transcription of raw audio has been processed using the language model neural network; and processing the first input using the language model neural network in response to determining that the specified time interval has elapsed, wherein the live transcription is a transcription of raw audio captured during the specified time interval (Zhu: ¶ 77, 207, 448, 462, 523-525: such as during a particular meeting by providing streaming speech, real time correction and summarization). Examiner takes official notice that periodic or scheduled summary of an ongoing or real time transcription was well known in the art before the effective filing date of the instant invention and would have comprised an obvious inclusion for at least the purpose of determining particular times for particular topics or relevance criteria such as a time duration of a speech, meeting, etc. The claim is thus considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 9
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 8, wherein either the first input prompt or the second prompt or both comprise an instruction to correct transcriptions generated from earlier live transcriptions of raw audio before the specified time interval (Zhu: ¶ 72, 82-84, 454, 541-544; Fig 9: system processes live speech input(s), such as by generation of speech to text, tagging, labelling, tokenization, etc. in concert with a language model, and other modalities, such as to make the recorded audio pliant to downstream processing such as by summarization, etc.); (Ma: Fig 2: system utilizes additional prompts to correct, edit, curate, etc. the summary of additional transcriptions). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 10
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 8, the operations further comprising: determining that transcribing has terminated; and, in response to determining that transcribing has terminated, processing a final input to generate text that summarizes the live transcription of raw audio data captured prior to, during, and after the specified time interval (Zhu: ¶ 72, 75, 82-84, 90-93; 454, 541-544; Fig 9: system processes live speech input(s), such as by generation of speech to text, tagging, labelling, tokenization, etc. in concert with a language model, and other modalities, such as to make the recorded audio pliant to downstream processing such as by summarization, etc.; such as at the end of a meeting, in receipt of an end tag or end keyword, etc.). Examiner takes official notice that periodic or scheduled summary of an ongoing or real time transcription was well known in the art before the effective filing date of the instant invention and would have comprised an obvious inclusion for at least the purpose of determining particular times for particular topics or relevance criteria such as a time duration of a speech, meeting, etc. The claim thus is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 11
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 1, wherein the plurality of text tokens comprise a set of speaker identifiers and a block of text associated with each speaker identifier (Zhu: ¶ 72, 82-84, 454, 489, 541-544; Fig 5, 9: system tokenizes inputs and ascribes same to particular identified speakers, roles thereof). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 12
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 1, wherein the prompt comprises a query to correct the live transcription and one or more additional instructions (Ma: Fig 2). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 13
Zhu in view of Zha teaches or suggests:
The user electronic device of claim 1, the operations further comprising outputting the modified transcription to a user of the user electronic device (Zhu: Fig 5); (Ma: Fig 2: a corrected transcription returned). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 14
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 1, wherein the first input prompt further comprises an instruction to identify action items for a particular speaker, or both, and the text output comprises text derived from the corrected transcription that identifies one or more action items for the particular speaker (Zhu: ¶ 65, 72, 82-84, 454, 489, 541-544; Fig 5, 9: speakers, roles, etc. identified, tasks ascribed thereto; in the case of the figure 5 illustration the transcriptions apply to product design and the summary one for the industrial designer speaker, role, etc.). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 15
Zhu in view of Ma teaches or suggests:
The user electronic device of claim 1, wherein the first input prompt further comprises an instruction to summarize the corrected transcript (Zhu: ¶ 462: such as by generating a meeting summary, etc.); (Ma: Abstract: such as additionally using the taught LLM for an additional prompt or additional item in a given prompt). The claim is considered obvious over Zhu as modified by Ma as addressed in the base claim as it would have been obvious to apply the further teaching of Zhu and/or Ma to the modified device of Zhu and Ma; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claims 16, 19—the claims are considered to recite substantially similar subject matter to that of claim 1 and are similarly rejected.
Regarding claims 17—the claim is considered to recite substantially similar subject matter to that of claim 3 and is similarly rejected.
Regarding claims 18—the claim is considered to recite substantially similar subject matter to that of claim 4 and is similarly rejected.
Regarding claim 20-22
Zhu in view of Ma teaches or suggests:
The user electronic device of claims 1, 16, 19 wherein the speech transcriber and the language model neural network both run on the user electronic device (Zhu: ¶ 39, 451, 544; system 110 comprises modules operable to transcribe, post process and correct data with respect to input content in concert with distributed services such as a speech service and wherein embodiments include the modules, services, etc. embodied upon a speech enabled smart device). While Zhu in view of Zha does not explicitly discuss the modules of the device embodied upon a singular user device Examiner considers such an embodiment obvious to try as at before the effective filing date of the instant invention there existed the recognized problem of providing functionality of a system at a server, at a user device, or in a distributed manner at a combination of the two; as such there existed a finite number of identified, predictable potential solutions to such an implementation; further one or ordinary skill in the art at the relevant time could be relied upon not only to recognize the utility of operating discrete functional modules in keeping with the finite solutions, but also to pursue potential solutions in this regard and with a reasonable expectation of success, predictable results, etc. and without undue experimentation; as such it would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to locate transcription module(s); correction module(s); LLM implementation(s) thereof, etc. at a user device or device local to a user; one of ordinary skill in the art would have expected only predictable results therefrom.
Response to Arguments
Applicant's arguments filed 6/4/26 have been fully considered Applicant’s arguments in concert with claim amendments, see Remarks and Claims, filed 6/4/26, with respect to the rejection(s) of claim(s) 1-3-5, 7-22 under 35 USC 103 over Zhu in view of Zhang, and Zhu in view of Zhang in view of Zhu have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Zhu and Ma.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL C MCCORD whose telephone number is (571)270-3701. The examiner can normally be reached 730-630 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAUL C MCCORD/Primary Examiner, Art Unit 2692