DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to a judicial exception, namely an abstract idea, without significantly more.
Claim 1 recites a computer-implemented method comprising receiving audio data, generating diarization results using a joint speech recognition and speaker diarization model, and processing the diarization results using a large language model conditioned on a diarization prompt to predict updated diarization results.
This is an abstract idea because the claim recites receiving information, analyzing information, organizing information, and modifying information, specifically the receipt of audio data, generation of diarization results, alignment of predicted terms with speaker tokens, and post-processing of those results using semantic interpretation and an LLM.
Under Step 2A, Prong One, the claim recites a judicial exception in the form of a mental process and related abstract information processing, because the claimed operations amount to evaluating and matching predicted terms with speaker tokens, identifying misalignment, and producing updated diarization results.
Under Step 2A, Prong Two, the claim does not integrate the abstract idea into a practical application. The additional elements, including “computer-implemented method executed on data processing hardware,” “joint speech recognition and speaker diarization model,” “large language model (LLM),” and “diarization prompt,” are recited at a high level of generality and merely implement the abstract idea using generic computing components and model terminology.
Under Step 2B, the claim elements, considered individually and as an ordered combination, do not amount to significantly more than the abstract idea itself. The claim does not recite a specific technological improvement to computer functionality or to speech recognition or diarization technology. Rather, the claim uses generic data processing hardware and model-based processing to achieve the abstract result of post-processing diarization output.
Accordingly, claim 1 is rejected under 35 U.S.C. § 101.
Claims 2-10 depend directly or indirectly from claim 1 and further limit the abstract information processing recited therein.
Claim 2 recites that the corresponding identity-specific speaker token does not reveal the particular identity of the respective one of the speakers that spoke the respective predicted term.
Claim 3 recites that the particular identity comprises a name or role of the respective one of the speakers that spoke the respective predicted term.
Claim 4 recites that processing the diarization results to predict updated diarization results comprises replacing the identity-agnostic speaker tokens with identity-specific speaker tokens.
Claim 5 recites identifying a predicted term misaligned with a corresponding identity-agnostic speaker token using semantic interpretation, realigning the identified predicted term with another one of the identity-agnostic speaker tokens, and generating the updated diarization results based on the realigned predicted term.
Claim 6 recites that the LLM is pre-trained on a diverse range of text data sourced from web documents, books, and code.
Claim 7 recites fine-tuning the LLM on training examples to perform post-processing on the diarization results.
Claim 8 recites that the diarization prompt comprises a single-shot learning example. Claim 9 recites that the single-shot learning example comprises an example input and output for conditioning the LLM.
Claim 10 recites that the diarization prompt comprises context data associated with the conversation.
These additional limitations do not add significantly more because they merely further specify the abstract information-processing steps using generic and conventional model features, training, prompting, and post-processing.
Accordingly, claims 2-10 are rejected under 35 U.S.C. § 101 for the same reasons as claim 1.
Claim 11 is directed to a system comprising data processing hardware and memory hardware storing instructions that when executed cause the data processing hardware to perform operations comprising receiving audio data, generating diarization results using a joint speech recognition and speaker diarization model, and processing the diarization results using an LLM conditioned on a diarization prompt to predict updated diarization results.
For the reasons set forth with respect to claim 1, claim 11 recites the abstract idea of receiving, analyzing, organizing, and modifying information, and the additional recited computer components are generic and do not integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself.
Accordingly, claim 11 is rejected under 35 U.S.C. § 101.
Claims 12-20 depend directly or indirectly from claim 11 and recite limitations corresponding to those in claims 2–10.
Claim 12 recites that the corresponding identity-specific speaker token does not reveal the particular identity of the respective one of the speakers that spoke the respective predicted term.
Claim 13 recites that the particular identity comprises a name or role of the respective one of the speakers that spoke the respective predicted term.
Claim 14 recites that processing the diarization results to predict updated diarization results comprises replacing the identity-agnostic speaker tokens with identity-specific speaker tokens.
Claim 15 recites identifying a predicted term misaligned with a corresponding identity-agnostic speaker token using semantic interpretation, realigning the identified predicted term with another one of the identity-agnostic speaker tokens, and generating the updated diarization results based on the realigned predicted term.
Claim 16 recites that the LLM is pre-trained on a diverse range of text data sourced from web documents, books, and code.
Claim 17 recites fine-tuning the LLM on training examples to perform post-processing on the diarization results.
Claim 18 recites that the diarization prompt comprises a single-shot learning example.
Claim 19 recites that the single-shot learning example comprises an example input and output for conditioning the LLM.
Claim 20 recites that the diarization prompt comprises context data associated with the conversation.
These limitations do not materially alter the abstract nature of the claim or provide significantly more than generic model training, prompting, and information-processing steps.
Accordingly, claims 12-20 are rejected under 35 U.S.C. § 101 for the same reasons as claim 11.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-4, 6-14, 16-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Park et al. (US 20250078842 A1).
Regarding claims 1 and 11, Park discloses a system comprising :data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations and a computer-implemented method executed on data processing hardware that causes the data processing hardware (i.e., “Audio processing server 102 may include a memory 104 (e.g., one or more memory devices or units) communicatively coupled with one or more processing devices”) (see paragraph 24) to perform operations comprising: receiving audio data comprising a plurality of spoken terms spoken by one or more speakers during a conversation (i.e., Audio processing server 102 may be configured to receive audio data 101 that may be associated with any speech episode involving one or more speakers.”…“Speech episodes may include a public or private conversation, a business meeting, a public or private presentation, an artistic event, a debate, an interaction between a digital agent (e.g., chatbot, digital avatar, etc.) and one or more users, an in-vehicle communication) (see paragraph 22); generating, using a joint speech recognition and speaker diarization model, diarization results based on the plurality of spoken terms spoken by the one or more speakers during the conversation, the diarization results comprising: a speech recognition result comprising a series of predicted terms (i.e., Multi-speaker speech recognition combines speaker diarization (SD) which maps various portions A.sub.1, A.sub.2, etc., of a given audio data to respective speakers S.sub.1, S.sub.2, etc. with automatic speech recognition (ASR)—which converts the audio portions A.sub.1, A.sub.2, etc., into spoken words W.sub.1, W.sub.2, etc.) (see paragraph 17) and a series of identity-agnostic speaker tokens, wherein each respective predicted term from the series of predicted terms is aligned with a corresponding identity- agnostic speaker token from the series of identity-agnostic speaker tokens and each corresponding identity-agnostic speaker token represents a generic identity of a respective one of the speakers that spoke the respective predicted term (i.e., “The SD branch identifies most likely speakers S* (whose number may be apriori unknown) responsible for uttering words captured by various audio portions A.”…“Training engine 162 may also train SD model 122 to associate specific portions (e.g., 0.05-5 sec portions) of training audio data 152 with various speakers, e.g., assigning unique labels to such portions, e.g., ‘Speaker 1,’ ‘Speaker 2,’ etc.”) (see paragraphs 17 and 27); and processing, using a large language model (LLM), the diarization results conditioned on a diarization prompt to predict, as output from the LLM, updated diarization results comprising (i.e., “The spoken words W.sub.1, W.sub.2, etc., and the set of speakers S.sub.1, S.sub.2, etc., may be used to form one or more prompts to the LM.”…“For example, a first prompt may inform the LM about a set of previously identified words {W.sub.p} and ask the LM to estimate likelihoods of various possible words W that follow this set {W.sub.p}: P(W|{W.sub.p}).”…“A second prompt may inform the LM about the most likely word W … and ask the LM to estimate likelihoods that various previously identified speakers S.sub.1, S.sub.2, etc. … have spoken the word W … : P(S|W).” (see paragraphs 19, 47, 54, and 65): the speech recognition result comprising the series of predicted terms and a series of identity-specific speaker tokens, wherein each respective predicted term from the series of predicted terms is aligned with a corresponding identity- specific speaker token from the series of identity-specific speaker tokens and each corresponding identity-specific speaker token representing a particular identity of a respective one of the speakers that spoke the respective predicted term (i.e., “The obtained word-to-speaker mapping W* .Math.S*.”…“the most probable path … indicates that the word ‘weekend’ … was spoken by Speaker A and … the word ‘yes’ … was spoken by Speaker B.” ) (see paragraphs 19 and 80).
Regarding claims 2 and 12, Park discloses a method and system as described in claims 1 and 11 above, wherein the corresponding identity-specific speaker token does not reveal the particular identity of the respective one of the speakers that spoke the respective predicted term (i.e., assigning unique labels to such portions, e.g., ‘Speaker 1,’ ‘Speaker 2,’ etc.) (see paragraph 27).
Regarding claims 3 and 13, Park discloses a method and system as described in claims 1 and 11 above, wherein the particular identity comprises a name or role of the respective one of the speakers that spoke the respective predicted term (i.e., “the most probable path … indicates that the word ‘weekend’ … was spoken by Speaker A and … the word ‘yes’ … was spoken by Speaker B.”) (see paragraph 80).
Regarding claims 4 and 14, Park discloses a method and system as described in claims 1 and 11 above, wherein processing the diarization results to predict updated diarization results comprises replacing the identity-agnostic speaker tokens with identity-specific speaker tokens (i.e., “The spoken words W.sub.1, W.sub.2, etc., and the set of speakers S.sub.1, S.sub.2, etc., may be used to form one or more prompts to the LM.”…“The obtained word-to-speaker mapping W* .Math.S*.”) (see paragraph 19).
Regarding claims 6 and 16, Park discloses a method and system as described in claims 1 and 11 above, wherein the LLM is pre-trained on a diverse range of text data sourced from web documents, books, and code (i.e., “Training data for language models includes many readily available texts in a practically unlimited number of different fields.”…“LM 124 may be further trained using training data containing a large number of texts, such as human dialogues, newspaper texts, magazine texts, book texts, web-based texts, and/or any other texts.”) (see paragraphs 18 and 30).
Regarding claims 7 and 17, Park discloses a method and system as described in claims 1 and 11 above, wherein the operations further comprise fine-tuning the LLM on training examples to perform post-processing on the diarization results (i.e., “LM 124 may be trained using training data containing a large number of texts”) (see paragraph 30).
Regarding claims 8 and 18, Park discloses a method and system as described in claims 1 and 11 above, wherein the diarization prompt comprises a single-shot learning example (i.e., “In some embodiments, LM 124 may be an N-gram language model, a large language model, or some other language model.”… “Multi-speaker speech recognition system 202 may use LM prompts 322 to generate requests to LM 124”) (see paragraphs 53-54. Also refer to paragraphs 29 and 30).
Regarding claims 9 and 19, Park discloses a method and system as described in claims 8 and 18 above, wherein the single-shot learning example comprises an example input and output for conditioning the LLM (i.e., training inputs and training/target outputs) (see paragraphs 18, 29).
Regarding claims 10 and 20, Park discloses a method and system as described in claims 1 and 11 above, wherein the diarization prompt comprises context data associated with the conversation (i.e., “a first prompt may inform the LM about a set of previously identified words {W.sub.p} and ask the LM to estimate likelihoods of various possible words W that follow this set {W.sub.p}: P(W|{W.sub.p}).”…“A second prompt may inform the LM about the most likely word W … and ask the LM to estimate likelihoods that various previously identified speakers S.sub.1, S.sub.2, etc. … have spoken the word W”) (see paragraphs 19). Also refer to paragraph 59 (i.e., “LM prompts 322 may generate a next-word prompt 324 that includes previously identified words {W.sub.p}”).
Allowable Subject Matter
Claims 5 and 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20240347064 A1, Li et al., Systems and Methods for Enhanced Speaker Diarization
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PIERRE LOUIS DESIR whose telephone number is (571)272-7799. The examiner can normally be reached Monday-Friday 9AM-5:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659