DETAILED ACTION
Introduction
This office action is in response to Applicant’s submission filed on March 18, 2025.
Claims 1-20 are pending in the application. As such, claims 1-20 have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings were received on March 18, 2025. These drawings have been accepted and considered by the Examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: the various “modules” claimed in claims 14 and 15.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. Specifically, each “module” is interpreted as being a “processor” as described in the specification [0008].
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1, 6 and 16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite:
An interactive method based on a large model, comprising:
acquiring a request speech;
performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics; and
processing the speech recognition feature using the large model to obtain a response text,
wherein the response text comprises a plurality of response words arranged in sequence,
a target response word among the plurality of response words is determined by processing the speech recognition feature and an associated response word feature using an attention fusion layer of the large model, and the associated response word feature is related to an associated response word arranged before the target response word.
The claim limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind. For example,
“acquiring a request speech” in the context of this claim encompasses a person listening to another person,
“performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics” in the context of this claim encompasses a person understanding the request,
“processing the speech recognition feature using the large model to obtain a response text” in the context of this claim encompasses a person providing an answer,
“a target response word among the plurality of response words is determined by processing the speech recognition feature and an associated response word feature using an attention fusion layer of the large model, and the associated response word feature is related to an associated response word arranged before the target response word” in the context of this claim encompasses a person ensuring the answer is thorough.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
The dependent claims do not add limitations that would either integrate the recited abstract idea into a practical application or could help the Claim as a whole to amount to significantly more than the Abstract idea identified for the Independent Claim.
Claims 2 and 17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite:
wherein the processing the speech recognition feature using the large model to obtain a response text comprises:
processing, using a text feature fusion layer, an initial associated response word feature representing the associated response word, so as to obtain the associated response word feature, wherein the large model comprises the text feature fusion layer;
processing the speech recognition feature and the associated response word feature using the attention fusion layer to obtain a target response word feature;
and determining the target response word based on the target response word feature.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“wherein the processing the speech recognition feature using the large model to obtain a response text comprises” in the context of this claim encompasses a person listening and providing a response,
“processing, using a text feature fusion layer, an initial associated response word feature representing the associated response word, so as to obtain the associated response word feature, wherein the large model comprises the text feature fusion layer” in the context of this claim encompasses a person ensuring to respond relevant to the question,
“processing the speech recognition feature and the associated response word feature using the attention fusion layer to obtain a target response word feature” in the context of this claim encompasses a person ensuring to respond relevant to the question,
“determining the target response word based on the target response word feature” in the context of this claim encompasses a person ensuring to respond relevant to the question.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the large model comprises N text feature fusion layers and N attention fusion layers, and N is an integer greater than 1;
wherein the processing the speech recognition feature and the associated response word feature using the attention fusion layer to obtain a target response word feature comprises:
processing, using an nth attention fusion layer, the speech recognition 1 feature and an nth associated response word feature, so as to obtain an nth intermediate fusion feature, wherein N>n>1, the n* associated response word feature is determined by processing an (n-1)th intermediate fusion feature using an nth text feature fusion layer, and a first associated response word feature is determined by processing the initial associated response word feature using a first text feature fusion layer; and
determining the target response word feature based on an Nth intermediate fusion feature in a case of n=N.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“the large model comprises N text feature fusion layers and N attention fusion layers, and N is an integer greater than 1” in the context of this claim encompasses a person ensuring to select a model with these basic parameters,
“processing the speech recognition feature and the associated response word feature using the attention fusion layer to obtain a target response word feature comprises” in the context of this claim encompasses a person listening and providing an answer,
“processing, using an nth attention fusion layer, the speech recognition 1 feature and an nth associated response word feature, so as to obtain an nth intermediate fusion feature, wherein N>n>1, the n* associated response word feature is determined by processing an (n-1)th intermediate fusion feature using an nth text feature fusion layer, and a first associated response word feature is determined by processing the initial associated response word feature using a first text feature fusion layer” in the context of this claim encompasses a person ensuring to select a model with these basic parameters,
“determining the target response word feature based on an Nth intermediate fusion feature in a case of n=N” in the context of this claim encompasses a person listening and providing an answer.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics comprises:
performing a feature extraction on the request speech to obtain an initial speech feature;
decoding the initial speech feature to obtain a plurality of initial decoding features, wherein the initial decoding feature represents a request word in the request speech; and
fusing the plurality of initial decoding features and the initial speech feature based on an attention mechanism to obtain the speech recognition feature.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics” in the context of this claim encompasses a person identifying the question properly,
“performing a feature extraction on the request speech to obtain an initial speech feature” in the context of this claim encompasses a person identifying a feature,
“decoding the initial speech feature to obtain a plurality of initial decoding features, wherein the initial decoding feature represents a request word in the request speech” in the context of this claim encompasses a person identifying characteristics of the feature,
“fusing the plurality of initial decoding features and the initial speech feature based on an attention mechanism to obtain the speech recognition feature” in the context of this claim encompasses a person understanding the question.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites no additional elements.
Accordingly, these no additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the no additional elements do not provide an inventive concept. The claim is not patent eligible.
Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the fusing the plurality of initial decoding features and the initial speech feature based on an attention mechanism to obtain the speech recognition feature comprises:
fusing the initial decoding feature and the initial speech feature based on the attention mechanism to obtain a request word audio feature corresponding to the request word;
performing a global feature fusion on a plurality of request word audio features to obtain an intermediate speech feature; and
fusing the intermediate speech feature and the plurality of initial decoding features based on the attention mechanism to obtain the speech recognition feature.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“fusing the plurality of initial decoding features and the initial speech feature based on an attention mechanism to obtain the speech recognition feature” in the context of this claim encompasses a person identifying the feature characteristics,
“fusing the initial decoding feature and the initial speech feature based on the attention mechanism to obtain a request word audio feature corresponding to the request word” in the context of this claim encompasses a person identifying the question properly,
“performing a global feature fusion on a plurality of request word audio features to obtain an intermediate speech feature” in the context of this claim encompasses a person identifying the feature properly,
“fusing the intermediate speech feature and the plurality of initial decoding features based on the attention mechanism to obtain the speech recognition feature” in the context of this claim encompasses a person identifying the feature properly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites no additional elements.
Accordingly, these no additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the no additional elements do not provide an inventive concept. The claim is not patent eligible.
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the initial large model is determined by training an extended large model, and
the extended large model is obtained by updating a network structure of a pre-trained basic large model based on an extended attention fusion layer.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“the initial large model is determined by training an extended large model” in the context of this claim encompasses a person selecting and using a model accordingly,
“the extended large model is obtained by updating a network structure of a pre-trained basic large model based on an extended attention fusion layer” in the context of this claim encompasses a person selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the basic large model comprises a multi-level basic feature fusion network, and
the basic feature fusion network comprises cascaded basic text feature fusion layer and basic feed-forward layer;
and wherein the extended large model comprises a multi-level extended feature fusion network, and
the extended feature fusion network comprises cascaded basic text feature fusion layer, extended attention fusion layer, and basic feed-forward layer.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“the basic large model comprises a multi-level basic feature fusion network” in the context of this claim encompasses a person selecting and using a model accordingly,
“the basic feature fusion network comprises cascaded basic text feature fusion layer and basic feed-forward layer” in the context of this claim encompasses a person selecting and using a model accordingly,
“the extended large model comprises a multi-level extended feature fusion network” in the context of this claim encompasses a person selecting and using a model accordingly,
“the extended feature fusion network comprises cascaded basic text feature fusion layer, extended attention fusion layer, and basic feed-forward layer” in the context of this claim encompasses a person selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model
a feature fusion network.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the initial large model is determined based on following training operations:
acquiring the extended large model, the pre-trained basic large model, a sample extension request text and a label extension response text;
processing the sample extension request text using the basic large model to obtain a sample extension request text feature;
processing the sample extension request text feature using the extended large model to obtain a sample extension response text; and
training the extended large model based on a difference between the sample extension response text and the label extension response text, so as to obtain the initial large model.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“acquiring the extended large model, the pre-trained basic large model, a sample extension request text and a label extension response text” in the context of this claim encompasses a person selecting and using a model accordingly,
“processing the sample extension request text using the basic large model to obtain a sample extension request text feature” in the context of this claim encompasses a person identifying a feature,
“processing the sample extension request text feature using the extended large model to obtain a sample extension response text” in the context of this claim encompasses a person providing an answer,
“training the extended large model based on a difference between the sample extension response text and the label extension response text, so as to obtain the initial large model” in the context of this claim encompasses a person selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the sample speech recognition feature is determined by processing the sample request speech using a speech recognition large model; and
wherein the training the initial large model based on a response text difference between the sample response text and the label response text to obtain a trained large model comprises:
determining a response text loss value based on the response text difference between the sample response text and the label response text; and
adjusting a model parameter of the speech recognition large model and amodel parameter of the initial large model based on the response text loss value and a request feature loss value to obtain the trained large model, wherein the request feature loss value represents a difference between the sample speech recognition feature and a preset label request text feature.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“the sample speech recognition feature is determined by processing the sample request speech using a speech recognition large model” in the context of this claim encompasses a person identifying a feature,
“determining a response text loss value based on the response text difference between the sample response text and the label response text” in the context of this claim encompasses a person deciding how accurate the answer measures,
“adjusting a model parameter of the speech recognition large model and amodel parameter of the initial large model based on the response text loss value and a request feature loss value to obtain the trained large model, wherein the request feature loss value represents a difference between the sample speech recognition feature and a preset label request text feature” in the context of this claim encompasses a person selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the label request text feature is determined by processing a sample request text using a preset large model, and the sample request speech is determined based on the sample request text.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“wherein the label request text feature is determined by processing a sample request text using a preset large model, and the sample request speech is determined based on the sample request text” in the context of this claim encompasses a person understanding the request and a feature.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein the sample request speech comprises sample request text audio information representing the sample request text, and sample environment audio information representing a speech environment sound.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“wherein the sample request speech comprises sample request text audio information representing the sample request text, and sample environment audio information representing a speech environment sound” in the context of this claim encompasses a person ensuring the specified data is present.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites no additional elements.
Accordingly, these no additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the no additional elements do not provide an inventive concept. The claim is not patent eligible.
Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
wherein a plurality of sample request speeches are provided, the plurality of sample request speeches have different sample speech attributes, and the sample speech attributes comprise one or more of: a timbre attribute, a speed attribute, a gender attribute, or an accent attribute.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“wherein a plurality of sample request speeches are provided, the plurality of sample request speeches have different sample speech attributes, and the sample speech attributes comprise one or more of: a timbre attribute, a speed attribute, a gender attribute, or an accent attribute” in the context of this claim encompasses a person ensuring the specified data is present.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites no additional elements.
Accordingly, these no additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the no additional elements do not provide an inventive concept. The claim is not patent eligible.
Claims 14 and 15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite:
an input module configured to receive input information;
a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and call the large model to implement themethod of claim 1, so as to obtain output information; and
an output module configured to output the output information obtained by the processing module.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and call the large model to implement themethod of claim 1, so as to obtain output information; and an output module configured to output the output information obtained by the processing module” in the context of this claim encompasses a person ensuring the specified data is present and selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
an input module
a processing module
an output module
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites:
an electronic device, comprising:
at least one processor; and
a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, are configured to cause the at least one processor to implement the method of claim 6.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, are configured to cause the at least one processor to implement the method of claim 6” in the context of this claim encompasses a person ensuring the specified data is present and selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
an electronic device
at least one processor
a memory
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claims 19 and 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite:
a non-transitory computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are configured to cause a computer to implement the method of claim 1.
The additional limitations of the claim do not preclude the method from practically being performed in the mind. For example,
“a non-transitory computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are configured to cause a computer to implement the method of claim 1” in the context of this claim encompasses a person ensuring the specified data is present and selecting and using a model accordingly.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites these additional elements. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea.
a non-transitory computer-readable storage medium
a large model.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea that do not provide an inventive concept. The claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6-7, 9-16 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu et al. (US Patent Pub. No. 20220230628 A1), hereinafter Zhu, in view of Wang et al. (US Patent Pub. No. 20220317641 A1), hereinafter Wang, in view of Mishima et al. (US Patent Pub. No. 20230252344 A1), hereinafter Mishima.
Regarding claims 1, 6 and 16, Zhu teaches an interactive method, a method of training a large model, and an electronic device (Zhu in [0008] teaches using a method, and in [0029] teaches training a model, and in [0030] teaches using a computing system)
based on a large model (Zhu in [0071] teaches using a multi-layer language module),
comprising:
[claim 16 only] at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, are configured to cause the at least one processor to at least (Zhu in [0030] teaches using a computing system with processors and storage for executable instructions):
acquiring a request speech (Zhu in [0071] teaches using a speech module which is configured to produce acoustic output embeddings based on input acoustic data);
performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics (Zhu in [0109] teaches using automatic speech recognition (ASR) machine learning models, and in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts);
and
processing the speech recognition feature using the large model to obtain a response text (Zhu in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts, and in [0061] teaches the computing system is able to operate the speech model to understand the acoustic data extracted from the electronic content, comprising speech, and by performing speech to text natural language processing and to generate text output based on the understanding of the extracted acoustic data),
wherein the response text comprises a plurality of response words arranged in sequence (Zhu in [0100] teaches the output is aligned through sequence-level alignment or token-level alignment).
[claim 6 only] training the initial large model based on a response text difference between the sample response text and the label response text to obtain a trained large model (Zhu in [0029] teaches using a computing system which is configured to train a plurality of machine learning models for speech recognition, natural language understanding, text-to-speech, and more particularly, training machine learning models to generate an integrated knowledge-language module, an integrated knowledge-speech module and/or an optimized speech model, and in [0119] teaches using a training data set comprising paired acoustic data and transcript data).
Zhu does not teach, however Wang teaches
an attention fusion layer (Wang in [0310] teaches using a multi-head attention layer of a fusion module).
Wang is considered to be analogous to the claimed invention because it is in the same field of attention layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu further in view of Wang to allow for using a multi-head attention layer of a fusion module. Motivation to do so would allow for by means of multi-modal data, intention classification, behavior pattern analysis, scene analysis, the user intention is recognized, the correct execution device is determined, execution logic is formed, and potential device conflicts and user conflicts are simultaneously solved (Wang [0291]).
Zhu, as modified above, teaches the plurality of response words, the speech recognition feature, the attention fusion layer, and the large model.
Zhu, as modified above, does not teach, however Mishima teaches
a target response word among the [plurality of response words] is determined by processing the [speech recognition feature] and an associated response word feature using an [attention fusion layer of the large model], and the associated response word feature is related to an associated response word arranged before the target response word (Mishima in [0041] teaches the answer text converter converts the fused feature into a plurality of answer word vectors respectively corresponding to a plurality of words constituting the predicted answer text).
Mishima is considered to be analogous to the claimed invention because it is in the same field of language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Mishima to allow for converting the fused feature into a plurality of answer word vectors. Motivation to do so would allow for a machine learning apparatus capable of learning a statistical model of a VQA task with high efficiency, a machine learning method, and an inference apparatus (Mishima [0106]).
Regarding claim 7, Zhu, as modified above, teaches the method according to claim 6.
Zhu, as modified above, teaches the large model.
Zhu further teaches
wherein the initial large model is determined by training an [extended] large model (Zhu in [0029] teaches using a computing system which is configured to train a plurality of machine learning models for speech recognition, natural language understanding, text-to-speech, and more particularly, training machine learning models to generate an integrated knowledge-language module, an integrated knowledge-speech module and/or an optimized speech model, and in [0119] teaches using a training data set comprising paired acoustic data and transcript data),
and
the extended large model is obtained by updating a network structure of a pre-trained basic large model [based on an extended attention fusion layer] (Zhu in [0029] teaches using a computing system which is configured to train a plurality of machine learning models for speech recognition, natural language understanding, text-to-speech, and more particularly, training machine learning models to generate an integrated knowledge-language module, an integrated knowledge-speech module and/or an optimized speech model, and in [0119] teaches using a training data set comprising paired acoustic data and transcript data, and in [0008] teaches the language module is pre-trained, and in [0040] teaches the neural model, which is configured as a neural network that is trainable or is trained to convert input text to speech data).
Zhu, as modified above, does not teach, however Wang teaches
an extended attention fusion layer (Wang in [0310] teaches using a multi-head attention layer of a fusion module, and in [0208] teaches the feature fusion network and the device selection network may perform training separately or jointly).
Wang is considered to be analogous to the claimed invention because it is in the same field of attention layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Wang to allow for using a multi-head attention layer of a fusion module. Motivation to do so would allow for by means of multi-modal data, intention classification, behavior pattern analysis, scene analysis, the user intention is recognized, the correct execution device is determined, execution logic is formed, and potential device conflicts and user conflicts are simultaneously solved (Wang [0291]).
Regarding claim 9, Zhu, as modified above, teaches the method according to claim 7.
Zhu, as modified above, teaches the initial large model, the extended large model, the basic large model, training an extended large model, and the pre-trained basic large model.
Zhu further teaches
wherein the initial large model is determined based on following training operations:
acquiring the extended large model, the pre-trained basic large model, a sample extension request text and a label extension response text (Zhu in [0029] teaches using a computing system which is configured to train a plurality of machine learning models for speech recognition, natural language understanding, text-to-speech, and more particularly, training machine learning models to generate an integrated knowledge-language module, an integrated knowledge-speech module and/or an optimized speech model, and in [0119] teaches using a training data set comprising paired acoustic data and transcript data [this may be repeated to obtain the sample extension request text and label extension response text]);
processing the sample extension request text using the basic large model to obtain a sample extension request text feature (Zhu in [0109] teaches using automatic speech recognition (ASR) machine learning models, and in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts [this may be repeated to obtain the sample extension request text feature]);
processing the sample extension request text feature using the extended large model to obtain a sample extension response text (Zhu in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts, and in [0061] teaches the computing system is able to operate the speech model to understand the acoustic data extracted from the electronic content, comprising speech, and by performing speech to text natural language processing and to generate text output based on the understanding of the extracted acoustic data [this may be repeated to obtain the sample extension response text]);
and
training the extended large model based on a difference between the sample extension response text and the label extension response text, so as to obtain the initial large model (Zhu in [0081] teaches the representation loss is minimized between the predicted token and the original token during training to optimize the language module).
Regarding claim 10, Zhu, as modified above, teaches the method according to claim 6.
Zhu, as modified above, teaches the sample request speech, the initial large model, the sample response text, the label response text, the trained large model, and the initial large model.
Zhu further teaches
wherein the sample speech recognition feature is determined by processing the sample request speech using a speech recognition large model (Zhu in [0109] teaches using automatic speech recognition (ASR) machine learning models, and in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts [this may be repeated to obtain the sample speech recognition feature]);
and
wherein the training the initial large model based on a response text difference between the sample response text and the label response text to obtain a trained large model comprises:
determining a response text loss value based on the response text difference between the sample response text and the label response text (Zhu in [0081] teaches the representation loss is minimized between the predicted token and the original token during training to optimize the language module [this may be repeated to obtain the response text loss value]);
and
adjusting a model parameter of the speech recognition large model and a model parameter of the initial large model based on the response text loss value and a [request feature loss value] to obtain the trained large model (Zhu in [0069] teaches using a technique referred to as joint-training, where in some instances, the integration or joint-training of the knowledge module and the language module enables the projection of the entities/relations included in the knowledge graph (e.g., knowledge graph data) and text (e.g., textual data) into a shared semantic latent space, and as the knowledge module produces representations from descriptive text, it solves the over-parameterization issue that can arise because entity embeddings are no longer part of the initial knowledge module's parameters),
wherein the request feature loss value represents a difference between the sample speech recognition feature and a preset label request text feature (Zhu in [0081] teaches the representation loss is minimized between the predicted token and the original token during training to optimize the language module [this may be repeated to obtain the request feature loss value]).
Regarding claim 11, Zhu, as modified above, teaches the method according to claim 10.
Zhu, as modified above, teaches the label request text feature, a sample request text, the large model, and the sample request speech.
Zhu further teaches
wherein the [label request text feature] is determined by processing a [sample request text] using a preset large model, and the sample request speech is determined based on the [sample request text] (Zhu in [0040] teaches converting input text to speech data using a neural network [here neural network maps to preset large model]).
Regarding claim 12, Zhu, as modified above, teaches the method according to claim 11.
Zhu, as modified above, teaches the label request text feature, a sample request text, the large model, and the sample request speech.
Zhu further teaches
wherein the sample request speech comprises sample request text audio information representing the sample request text, and sample environment audio information representing a speech environment sound (Zhu in [0040] teaches using acoustic data which comprises electronic content/data obtained from a speaker or a source comprising one or more speakers or a source comprising one or more speakers, background noise, non-human speakers and/or machine speakers).
Regarding claim 13, Zhu, as modified above, teaches the method according to claim 11.
Zhu, as modified above, teaches the request speech, and the sample speech.
Zhu further teaches
wherein a plurality of sample request speeches are provided, the plurality of sample request speeches have different sample speech attributes (Zhu in [0092] teaches speech utterances included in input acoustic data such as their phonetic content and speaker characteristics, and in some embodiments, the input to the speech module comprises a plurality of audio features, and in some instances, the audio features are based on 80-dimensional log Mel spectrograms),
and
the sample speech attributes comprise one or more of: a timbre attribute, a speed attribute, a gender attribute, or an accent attribute (Zhu in [0191] teaches using attributes about a personalized voice, including native language, secondary languages, user gender, voice prosody qualities, voice timbre qualities).
Regarding claims 14 and 15, Zhu, as modified above, teaches the method of claims 1 and 6.
Zhu further teaches
an intelligent agent (Zhu in [0035] teaches using personal assistant devices),
comprising:
an input module configured to receive input information (Zhu in [0071] teaches using a speech module which is configured to produce acoustic output embeddings based on input acoustic data);
a processing module configured to determine a target task based on the input information received by the input module (Zhu in [0059] teaches implement and operate an optimized speech model and/or an integrated knowledge-speech module and/or an integrated knowledge-language module to perform one or more natural language understanding tasks),
determine a [claim 15 only: initial] large model based on the target task (Zhu in [0059] teaches implement and operate an optimized speech model and/or an integrated knowledge-speech module and/or an integrated knowledge-language module to perform one or more natural language understanding tasks),
and
call the [claim 15 only: initial] large model to implement [the method of claim 1/claim 6], so as to obtain output information (Zhu in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts, and in [0061] teaches the computing system is able to operate the speech model to understand the acoustic data extracted from the electronic content, comprising speech, and by performing speech to text natural language processing and to generate text output based on the understanding of the extracted acoustic data);
and
an output module configured to output the output information obtained by the processing module (Zhu in [0030] teaches the computing system is also shown including user interface(s) and input/output (I/O) device(s)).
Regarding claim 18, Zhu, as modified above, teaches the method according to claim 6.
Zhu further teaches
an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, are configured to cause the at least one processor to implement the [method of claim 6] (Zhu in [0030] teaches using a computing system with processors and storage for executable instructions).
Regarding claims 19 and 20, Zhu, as modified above, teaches the method of claims 1 and 6.
Zhu further teaches
a non-transitory computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are configured to cause a computer to implement the [method of claim 1/claim 6] (Zhu in [0030] teaches using a computing system with processors and storage for executable instructions).
Claims 2, 3 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Wang, in view of Mishima, in view of Piergiovanni et al. (US Patent Pub. No. 20230394306 A1), hereinafter Piergiovanni.
Regarding claims 2 and 17, Zhu, as modified above, teaches the method and electronic device according to claims 1 and 16.
Zhu, as modified above, teaches the speech recognition feature, the response text, the associated response word feature, the large model, the associated response word, the associated response word feature, the speech recognition feature, the attention fusion layer, and the target response word.
Zhu, as modified above, does not teach, however Mishima teaches
wherein the processing the speech recognition feature using the large model to obtain a response text comprises:
[claim 17 only] wherein the instructions are further configured to cause the at least one processor to at least:
processing, [using a text feature fusion layer], an initial associated response word feature representing [the associated response word], so as to obtain the [associated response word feature], wherein the [large model] comprises the [text feature fusion layer] (Mishima in [0041] teaches the answer text converter converts the fused feature into a plurality of answer word vectors respectively corresponding to a plurality of words constituting the predicted answer text [this may be repeated to obtain the initial associated response word feature]);
processing the [speech recognition feature] and the [associated response word feature] using the [attention fusion layer] to obtain a target response word feature (Mishima in [0041] teaches the answer text converter converts the fused feature into a plurality of answer word vectors respectively corresponding to a plurality of words constituting the predicted answer text [this may be repeated to obtain the target response word feature]);
and determining the target response word based on the target response word feature (Mishima in [0041] teaches the answer text converter converts the fused feature into a plurality of answer word vectors respectively corresponding to a plurality of words constituting the predicted answer text [this may be repeated to obtain the target response word]).
Mishima is considered to be analogous to the claimed invention because it is in the same field of language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Mishima to allow for converting the fused feature into a plurality of answer word vectors. Motivation to do so would allow for a machine learning apparatus capable of learning a statistical model of a VQA task with high efficiency, a machine learning method, and an inference apparatus (Mishima [0106]).
Zhu, as modified above, does not teach, however Piergiovanni teaches
using a text feature fusion layer (Piergiovanni in [0065] the fusion layer reshapes the text features).
Piergiovanni is considered to be analogous to the claimed invention because it is in the same field of text feature fusion layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Piergiovanni to allow for a fusion layer to reshape the text features. Motivation to do so would allow for providing high quality performance in a multi-task setting, working simultaneously on many different tasks, without fine-tuning on individual tasks (Piergiovanni [0041]).
Regarding claim 3, Zhu, as modified above, teaches the method according to claim 2.
Zhu, as modified above, teaches the text feature fusion layer, the attention fusion layer, the speech recognition feature, the associated response word feature, and the target response word feature.
Zhu, as modified above, does not teach, however Piergiovanni teaches
wherein the large model comprises N text feature fusion layers and N attention fusion layers, and N is an integer greater than 1 (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results);
wherein the processing the speech recognition feature and the associated response word feature using the attention fusion layer to obtain a target response word feature comprises:
processing, using an nth attention fusion layer, the speech recognition feature and an nth associated response word feature, so as to obtain an nth intermediate fusion feature (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results [here the nth layers can be used to produce the nth features]),
wherein N>n>1, the n* associated response word feature is determined by processing an (n-1)th intermediate fusion feature using an nth text feature fusion layer (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results [here the nth layers can again be used to produce the nth features]),
and a first associated response word feature is determined by processing the initial associated response word feature using a first text feature fusion layer (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results [here the nth layers can yet again be used to produce the nth features]);
and
determining the target response word feature based on an Nth intermediate fusion feature in a case of n=N (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results [here the process can be repeated n times until n=N]).
Piergiovanni is considered to be analogous to the claimed invention because it is in the same field of text feature fusion layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Piergiovanni to allow for using a number (N) of fusion layers. Motivation to do so would allow for providing high quality performance in a multi-task setting, working simultaneously on many different tasks, without fine-tuning on individual tasks (Piergiovanni [0041]).
Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Wang, in view of Mishima, in view of Jeong et al. (US Patent Pub. No. 20070185712 A1), hereinafter Jeong.
Regarding claim 4, Zhu, as modified above, teaches the method according to claim 1.
Zhu, as modified above, teaches the speech recognition, the request speech, the speech recognition feature, and the request semantics.
Zhu further teaches
wherein the performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics comprises:
obtain an initial speech feature (Zhu in [0109] teaches using automatic speech recognition (ASR) machine learning models, and in [0060] teaches performing semantic analysis to understand semantic information from text-based transcripts [this can be repeated to obtain an initial speech feature]).
Zhu, as modified above, does not teach, however Wang teaches
fusing the [plurality of initial decoding features] and the [initial speech feature] based on an attention mechanism to obtain the [speech recognition feature] (Wang in [0310] teaches using a multi-head attention layer of a fusion module).
Wang is considered to be analogous to the claimed invention because it is in the same field of attention layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Wang to allow for using a multi-head attention layer of a fusion module. Motivation to do so would allow for by means of multi-modal data, intention classification, behavior pattern analysis, scene analysis, the user intention is recognized, the correct execution device is determined, execution logic is formed, and potential device conflicts and user conflicts are simultaneously solved (Wang [0291]).
Zhu, as modified above, does not teach, however Jeong teaches
performing a feature extraction on the [request speech] to obtain an [initial speech feature] (Jeong in [0045] teaches features are extracted from the input speech signal);
decoding the initial speech feature to obtain a plurality of initial decoding features, wherein the initial decoding feature represents a request word in the [request speech] (Jeong in [0045] teaches features are extracted from the input speech signal and the most similar feature to the decoded speech feature from words stored in a recognition list is recognized).
Jeong is considered to be analogous to the claimed invention because it is in the same field of decoding speech features. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Jeong to allow for features are extracted from the input speech signal. Motivation to do so would allow for measuring confidence of speech recognition by comparing a phase change point of a speech signal input to a speech recognizer and a phoneme string change point of a result of speech recognition and using the difference between the phase change point and the phoneme string change point, and a likelihood ratio (Jeong [0010]).
Regarding claim 5, Zhu, as modified above, teaches the method according to claim 4.
Zhu, as modified above, teaches the fusing, the plurality of initial decoding features, the initial speech feature, the speech recognition feature, and the initial decoding feature.
Zhu, as modified above, does not teach, however Wang teaches
wherein the fusing the plurality of initial decoding features and the initial speech feature based on an attention mechanism to obtain the speech recognition feature comprises:
fusing the [initial decoding feature] and the [initial speech feature] based on the attention mechanism to obtain a request word audio feature corresponding to the request word (Wang in [0310] teaches using a multi-head attention layer of a fusion module [this may be repeated to obtain the request word audio feature]);
performing a global feature fusion on a [plurality of request word audio features] to obtain an intermediate speech feature (Wang in [0310] teaches using a multi-head attention layer of a fusion module [this may be repeated to obtain the intermediate speech feature]);
and
fusing the [intermediate speech feature] and the [plurality of initial decoding features] based on the attention mechanism to obtain the speech recognition feature (Wang in [0310] teaches using a multi-head attention layer of a fusion module [this may be repeated to obtain the speech recognition feature]).
Wang is considered to be analogous to the claimed invention because it is in the same field of attention layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Wang to allow for using a multi-head attention layer of a fusion module. Motivation to do so would allow for by means of multi-modal data, intention classification, behavior pattern analysis, scene analysis, the user intention is recognized, the correct execution device is determined, execution logic is formed, and potential device conflicts and user conflicts are simultaneously solved (Wang [0291]).
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Zhu, in view of Wang, in view of Mishima, in view of Piergiovanni, in view of Shu et al. (US Patent Pub. No. 20230281785 A1), hereinafter Shu.
Regarding claim 8, Zhu, as modified above, teaches the method according to claim 7.
Zhu, as modified above, teaches the basic large model, and the extended large model.
Zhu, as modified above, does not teach, however Wang teaches
wherein the [basic large model] comprises a [multi-level] basic feature fusion network (Wang in [0310] teaches using a multi-head attention layer of a fusion module, and in [0208] teaches the feature fusion network and the device selection network may perform training separately or jointly),
and
wherein the [extended large model] comprises a [multi-level] extended feature fusion network (Wang in [0310] teaches using a multi-head attention layer of a fusion module, and in [0208] teaches the feature fusion network and the device selection network may perform training separately or jointly),
and
extended attention fusion layer (Wang in [0310] teaches using a multi-head attention layer of a fusion module, and in [0208] teaches the feature fusion network and the device selection network may perform training separately or jointly).
Wang is considered to be analogous to the claimed invention because it is in the same field of attention layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Wang to allow for using a multi-head attention layer of a fusion module. Motivation to do so would allow for by means of multi-modal data, intention classification, behavior pattern analysis, scene analysis, the user intention is recognized, the correct execution device is determined, execution logic is formed, and potential device conflicts and user conflicts are simultaneously solved (Wang [0291]).
Zhu, as modified above, does not teach, however Piergiovanni teaches
the basic [feature fusion network] comprises [cascaded] basic text feature fusion layer and basic feed-forward layer (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results [here the nth layers can yet again be used to produce the nth features], and in [0071] teaches using neural networks which can include feed-forward neural networks);
the extended feature fusion network comprises
[cascaded] basic text feature fusion layer (Piergiovanni in [0065] a number (N) of fusion layers can be applied to produce a set of results [here the nth layers can yet again be used to produce the nth features]),
and
basic feed-forward layer (Piergiovanni in [0071] teaches using neural networks which can include feed-forward neural networks).
Piergiovanni is considered to be analogous to the claimed invention because it is in the same field of text feature fusion layers. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Piergiovanni to allow for using a number (N) of fusion layers. Motivation to do so would allow for providing high quality performance in a multi-task setting, working simultaneously on many different tasks, without fine-tuning on individual tasks (Piergiovanni [0041]).
Zhu, as modified above, does not teach, however Shu teaches
a multi-level [basic feature fusion] network (Shu in [0010] teaches using a multi-level feature extraction instance segmentation network is obtained by cascading three levels of instance segmentation networks)
cascaded [basic text feature fusion] layer (Shu in [0010] teaches using a multi-level feature extraction instance segmentation network is obtained by cascading three levels of instance segmentation networks)
Shu is considered to be analogous to the claimed invention because it is in the same field of cascaded multi-level networks. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Zhu, as modified above, further in view of Shu to allow for using a cascaded multi-level network. Motivation to do so would allow for a method and system for defect detection, so as to detect defects when an object to be detected is very small, significantly reduce detection costs, and greatly improve quality detection efficiency (Shu [0005]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL J. MUELLER whose telephone number is (571)272-1875. The examiner can normally be reached M-F 9:00am-5:00pm (Eastern).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
PAUL MUELLER
Examiner
Art Unit 2657
/PAUL J. MUELLER/Examiner, Art Unit 2657