Prosecution Insights
Last updated: October 02, 2026
Application No. 18/469,909

RE-ARRANGING FEED FORWARD NETWORKS (FFNs) IN TRANSFORMER-BASED MODELS

Final Rejection §103§112
Filed
Sep 19, 2023
Examiner
WOOLWINE, SHANE D
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
2 (Final)
86%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
332 granted / 384 resolved
+31.5% vs TC avg
Strong +21% interview lift
Without
With
+20.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
9 currently pending
Career history
393
Total Applications
across all art units

Statute-Specific Performance

§101
13.3%
-26.7% vs TC avg
§103
49.9%
+9.9% vs TC avg
§102
16.1%
-23.9% vs TC avg
§112
12.2%
-27.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 384 resolved cases

Office Action

§103 §112
CTNF 18/469,909 CTNF 90409 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. 07-30-03-h AIA Claim Interpretation 07-30-03 AIA The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. 07-30-05 The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Such claim limitation(s) is/are: “An apparatus comprising: means for receiving, by an artificial neural network (ANN) model…” in claim 24 which is interpreted as being implemented using a processor as described in paragraph [0006] of the instant application. “means for processing, by a token interaction block of the ANN model…” in claim 24 which is interpreted as being implemented using a processor as described in paragraph [0006] of the instant application. “means for generating, by a feed forward network (FFN) block of the ANN model, a mixture of channel features…” in claim 24 which is interpreted as being implemented using a processor as described in paragraph [0006] of the instant application. “means for determining, by an attention block of the ANN model, a set of attended features of the mixture of channel features…” in claim 24 which is interpreted as being implemented using a processor as described in paragraph [0006] of the instant application. “means for generating, by the ANN model, an inference…” in claim 24 which is interpreted as being implemented using a processor as described in paragraph [0006] of the instant application. “means for processing the set of tokens…” in claim 25 which is interpreted as being implemented using a processor as described in paragraph [0032] of the instant application. “comprising means for processing the spatial mixture of the set of features…” in claim 26 which is interpreted as implemented using a processor as described in paragraph [0032] of the instant application. “means for dividing the image into a set of patches…” in claim 29 which is interpreted as implemented using a processor as described in paragraphs [0069]-[0070] of the instant application. “means for reshaping the mixture of channel features…” in claim 30 which is interpreted as implemented using a processor as described in paragraph [0073] of the instant application. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 6, 14, 22 and 29 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claims 6, 14, 22, and 29, taking claim 14 as exemplary: Claims 6, 14, 22, and 29 all recite “divide the image into a set of patches, the set of tokens corresponding to the set of patches of the image. ” However, claims 5, 13, 21, and 29, upon which claims 6, 14, 22, and 29 recite that an input contains “an image” as an alternative statement along with the possibility of including an audio signal or a textural input in place of “an image.” It is unclear how “the image” could be divided if it is not required as part of the proceeding claim limitation. Therefore, claims 6 , 14, 22, and 29 are rejected as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. For the purposes of this office action, claims 6, 14, 22, and 29 is interpreted as including a possible image within the input in coordination with claims 5, 13, 21, and 28. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-103 AIA The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. 07-20-02-aia AIA This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 07-21-aia AIA Claim (s) 1-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al., (US 2022/0277728 A1, hereinafter Zhang ) in view of YUAN K. et al., ("Incorporating Convolution Designs into Visual Transformers", 2021 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, 10 October 2021, pp. 559-568., part of the applicant submitted prior art and hereinafter Yuan ) . Regarding claims 1, 9, 17, and 24, taking claim 9 as exemplary: Zhang shows: “An apparatus, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: receive, by an artificial neural network (ANN) model, a set of tokens corresponding to an input;” ( Paragraph [0076]: “The exemplary neural TTS systems according to the embodiments of the disclosure are described above with reference to FIGS. 2, 11, and 12, and the exemplary methods for generating speech through neural TTS synthesis according to the embodiments of the present disclosure are described accordingly. The above systems and methods intend to generate speech corresponding to a text input based on both a phone feature and context features of the text input.” In paragraph [0079]: “At step 1320, a phone sequence may be identified from the text input, such as a current sentence, by various techniques such as LTS.” In paragraph [0080]: “At step 1330, the phone sequence may be updated by adding a begin token and/or an end token to the beginning and/or the end of the phone sequence , respectively. In an implementation, a mute phone may be used as the begin token and the end token .” In paragraph [0081]: “At step 1340, a phone feature may be generated based on the updated phone sequence. For example, the updated phone sequence may be converted into a phone embedding vector sequence through a phone embedding model. The phone embedding vector sequence may then be provided to an encoder, such as the encoder 110 in FIG. 1, to generate the phone feature corresponding to the updated phone sequence.” In paragraph [0133]: “The apparatus 1800 may comprise at least one processor 1810. The apparatus 1800 may further comprise a memory 1820 connected with the processor 1810. The memory 1820 may store computer-executable instructions that, when executed, cause the processor 1810 to perform any operations of the methods for generating speech through neural TTS synthesis according to the embodiments of the present disclosure as mentioned above.” And in paragraph [0134]: The embodiments of the present disclosure may be embodied in a non-transitory computer-readable medium . The non-transitory computer-readable medium may comprise instructions that, when executed, cause one or more processors to perform any operations of the methods for generating speech through neural TTS synthesis according to the embodiments of the present disclosure as mentioned above.” – the above processor and memory addresses the processor implemented method of claim 1 (dependent claims 2-8), the apparatus including a memory and processor of claim 9 (dependent claims 10-16), the non-transitory computer-readable medium having program code recorded thereon of claim 17 (dependent claims 18-23), and the apparatus means for that includes a processor (according the 112f analysis above) of claim 24 (dependent claims 25-30).) “process, by a token interaction block of the ANN model, the set of tokens according to each channel of the input to generate a spatial mixture of a set of features for each channel of the input;” ( Paragraph [0084]: “as shown in FIG. 11, the phone feature and the context features may be combined into first mixed features, and the first mixed features may be provided to an attention unit to obtain first attended mixed features. At the same time, global features in the context features may be cascaded with the first attended mixed features after passing its own attention unit, thereby obtaining second mixed features. The second mixed features may be provided to a decoder to generate acoustic features corresponding to the second mixed features. The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” In paragraph [0085]: “In yet another implementation, as shown in FIG. 12, the phone feature and the context features may be combined into first mixed featu res, and the first mixed features may be provided to the attention unit to obtain first attended mixed features. Meanwhile, averaging pooling may be performed on global features in the context features to obtain average global features. The average global features may be cascaded with the first attended mixed features to obtain second mixed features. The second mixed features may be provided to a decoder to generate acoustic features corresponding to the second mixed features. The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .”) “generate, by a feed forward network (FFN) block of the ANN model, a mixture of channel features based on the spatial mixture of the set of features for each channel of the input;” ( Paragraph [0036]: “The decoder 140 may include a pre-net 142 consisted of feed-forward layers, Long Short Term Memories (LSTMs) 144, a linear projection 146, a post-net 148 consisted of convolution layers , etc. The LSTMs 144 may receive an input from the pre-net 142 and provide its output to the linear projection 146 while the processing by the LSTMs 144 is affected by the attention unit 130. The linear projection 146 may provide its output to the pre-net 142 and the post-net 148, respectively. Finally, the output of the post-net 148 is combined with the output of the linear projection 146 to produce the acoustic features 150. In an implementation, the linear projection 146 may also be used to generate stop tokens.” And in paragraph [0041]: “The generated phone feature 210 and context features 218 may be combined into mixed features through a cascading unit 220. An attention unit 222 may apply an attention mechanism on the mixed features , such as a location sensitive attention mechanism. The attended mixed features may be provided to a decoder 224. The decoder 224 may correspond to the decoder 140 in FIG. 1. The decoder 224 may generate, based on the attended mixed features, acoustic features corresponding to the mixed features. The acoustic features may then be provided to a vocoder 226 . The vocoder 226 may correspond to the vocoder 160 in FIG. 1. A speech waveform 170 corresponding to the text input 202 may be generated by the vocoder 226 .” – The use of different layers to get input and the use of a cascading unit is the use of channels.) “and generate, by the ANN model, an inference based on the set of attend features of the mixture of channel features.” ( Paragraph [0082]: “At step 1350, a speech waveform corresponding to the text input may be generated based on the phone feature and context features . The speech waveform may be generated in a variety of ways .” In paragraph [0083]: “In an implementation, as shown in FIG. 2, the phone feature and the context features may be combined into mixed features, and the mixed features may be provided successively to an attention unit and a decoder to generate acoustic features corresponding to the mixed features. The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” In paragraph [0084]: “In another implementation, as shown in FIG. 11, the phone feature and the context features may be combined into first mixed features, and the first mixed features may be provided to an attention unit to obtain first attended mixed features ... The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” And in paragraph [0085]: “In yet another implementation, as shown in FIG. 12, the phone feature and the context features may be combined into first mixed feature s, and the first mixed features may be provided to the attention unit to obtain first attended mixed features... The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input. ”) But Zhang does not appear to explicitly recite “determine, by an attention block of the ANN model, a set of attended features of the mixture of channel features according to a set of attention weights;” However, Yuan teaches “determine, by an attention block of the ANN model, a set of attended features of the mixture of channel features according to a set of attention weights;” ( Page 561, column 3, paragraph 3: “MSA. For a self-attention (SA) modul e, the sequence of input tokens xt ∈ R(N+1)× C are linear transformed into qkv spaces, i.e., queries Q ∈ R(N+1)×C, keys K ∈ R(N+1)×C and values V ∈ R(N+1)×C. Then a weighted sum over all values in the sequence is computed through: Attention(Q,K,V) = softmax(QKT√C )V (2) And a linear transformation is performed to the weighted values .”) Zhang and Yuan are analogous in the arts because both Zhang and Yuan describe models with feature mixing. Therefore, it would be obvious to one of ordinary skill in the art at the filing date of the instant application, having the teachings of Zhang and Yuan before him or her, to modify the teachings of Zhang to include the teachings of Yuan in order to support the use of images as a feature input ( see Yuan page 561, column 2, paragraph 7 ) and thereby increase marketability of Zhang . Regarding claims 2, 10, 18, and 25, taking claim 10 as exemplary: Zhang and Yuan teach the method, apparatus, non-transitory computer-readable medium, and apparatus of claims 1, 9, 17, and 24 as claimed and specified above. And Yuan teaches “in which the at least one processor is further configured to process the set of tokens corresponding to the input using a depth-wise convolution to generate the spatial mixture of the set of features.” ( Page 562, Paragraph 1: "Figure 3: Illustration of the Locally-enhanced Feed-Forward module. First, patch tokens are projected into a higher dimension . Second, they are restored to “images” in the spatial dimension based on the original positions . Third, a depth-wise convolution is performed on the restored tokens as shown in the yellow region. Then the patch tokens are flattened and projected to the initial dimension. Besides, the class token conducts an identical mapping.") Regarding claims 3, 11, 19, and 26, taking claim 11 as exemplary: Zhang and Yuan teach the method, apparatus, non-transitory computer-readable medium, and apparatus of claims 1, 9, 17, and 24 as claimed and specified above. And Yuan teaches “in which the at least one processor is further configured to process the spatial mixture of the set of features using a point-wise convolution to generate the mixture of channel features.” ( Page 561, column 2, paragraphs 6-7: "FFN. FFNperforms point-wise operations , which are applied to each token separately. It consists of two linear transformations with a non-linear activation in between: FFN(x) = σ(xW1 +b1)W2 +b2 (3) where W1 ∈ RC× K is the weight of the first layer, projecting each token into a higher dimension K. And W2 ∈ RK×C is the weight of the second layer. b1 ∈ RK and b2 ∈ RC are the biases. And σ(·) is the non-linear activation of GELU [13] in ViT. Complementary to the MSA module, the FFN module performs dimensional expansion/reduction and non-linear transformation on each token , thereby enhancing the representation ability of tokens. However, the spatial relationship among tokens, which is important in vision, is not considered. This leads that the original ViT needs a mass of training data to learn these inductive biases.") Regarding claims 4, 12, 20, and 27, taking claim 12 as exemplary: Zhang and Yuan teach the method, apparatus, non-transitory computer-readable medium, and apparatus of claims 1, 9, 17, and 24 as claimed and specified above. And Zhang shows “in which the ANN model comprises a vision transformer, a bi-directional encoder representations from transformers (BERT), a robustly optimized BERT approach (RoBERTa)-based transformer, an XLNet-based transformer, Transformer-XL-based transformer, or a generative pre-trained transformer (GPT).” ( Paragraph [0048]: “The word embedding model 412 may be based on Natural Language Processing (NLP) techniques, such as Neural Machine Translation (NMT). Both the word embedding model and the neural TTS system have similar sequence-to-sequence encoder-decoder frameworks, this would benefit to network convergence. In one embodiment, a Bidirectional Encoder Representations from Transformers (BERT) model may be employed as the word embedding model. A word embedding vector sequence may be generated based on the word sequence 404 through the word embedding model 412, wherein each word has a corresponding embedding vector, and all of these embedding vectors form the word embedding vector sequence. The word embedding vector contains meaning and semantic context information of a word, which will facilitate improvement of naturalness of a generated speech.” – Note that the claim is written in the alternative and not all claim elements (i.e. a robustly optimized BERT approach (RoBERTa)-based transformer, an XLNet-based transformer, Transformer-XL-based transformer, or a generative pre-trained transformer (GPT)) must be recited for teaching by the reference to be satisfied.) Regarding claims 5, 13, 21, and 28, taking claim 13 as exemplary: Zhang and Yuan teach the method, apparatus, non-transitory computer-readable medium, and apparatus of claims 1, 9, 17, and 24 as claimed and specified above. And Zhang shows “in which the input comprises an image, an audio signal, or a textual input.” ( Paragraph [0082]: “At step 1350, a speech waveform corresponding to the text input may be generated based on the phone feature and context features . The speech waveform may be generated in a variety of ways.” And in paragraph [0083]: “In an implementation, as shown in FIG. 2, the phone feature and the context features may be combined into mixed features, and the mixed features may be provided successively to an attention unit and a decoder to generate acoustic features corresponding to the mixed features . The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” In paragraph [0084]: “In another implementation, as shown in FIG. 11, the phone feature and the context features may be combined into first mixed features, and the first mixed features may be provided to an attention unit to obtain first attended mixed features... The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” In paragraph [0085]: “In yet another implementation, as shown in FIG. 12, the phone feature and the context features may be combined into first mixed features, and the first mixed features may be provided to the attention unit to obtain first attended mixed features... The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” – Note that the claim is written in the alternative and not all claim elements (i.e. image input) needs to be recited for teaching by the reference to be satisfied.) Regarding claims 6, 14, 22, and 29, taking claim 14 as exemplary: Zhang and Yuan teach the method, apparatus, non-transitory computer-readable medium, and apparatus of claims 5, 13, 21, and 28 as claimed and specified above. And Zhang shows “in which the at least one processor is further configured to divide ... into a set of patches, the set of tokens corresponding to the set of patches of…” ( Paragraph [0086]: “Since mute phones are added as the begin token and/or the end token at the beginning and/or the end of the phone sequence of the text input , respectively, the generated speech waveform corresponding to the text input has a period of mute at the beginning and/or the end, respectively . Since the context features are taken into account when generating the speech waveform, the mute at the beginning and/or the end of the speech waveform is related to the context features, which may vary with the context features . For example, the mute at the end of the speech waveform of the text input and the mute at the beginning of a speech waveform of a next sentence of the text input constitute a pause between the text input and the next sentence. The duration of the pause may accordingly vary with the context features , such that the rhythm of the generated speech corresponding to a set of sentences is more abundant and natural.” – the use of periods is the dividing into patches.) But Zhang does not appear to explicitly recite the use of “images.” However, Yuan teaches the use of “images” ( Page 561, column 2, paragraph 7: "To solve the above-mentioned problems in tokenization, we propose a simple but effective module named as Image to-Tokens (I2T) that extracts patches from feature maps instead of raw input images." – the patches from feature maps associated with tokens is the use of dividing into patches associated with images.) Regarding claims 7 and 15, taking claim 15 as exemplary: Zhang and Yuan teach the method and apparatus of claims 1 and 9 above. And Zhang shows “in which the inference is a classification of the input.” ( Paragraph [0079]: “step 1320, a phone sequence may be identified from the text input, such as a current sentence, by various techniques such as LTS.” In paragraph [0080]: “At step 1330, the phone sequence may be updated by adding a begin token and/or an end token to the beginning and/or the end of the phone sequence, respectively. In an implementation, a mute phone may be used as the begin token and the end token.” In paragraph [0081]: “At step 1340, a phone feature may be generated based on the updated phone sequence. For example, the updated phone sequence may be converted into a phone embedding vector sequence through a phone embedding model. The phone embedding vector sequence may then be provided to an encoder, such as the encoder 110 in FIG. 1, to generate the phone feature corresponding to the updated phone sequence .” And in paragraph [0082]: “At step 1350, a speech waveform corresponding to the text input may be generated based on the phone feature and context features. The speech waveform may be generated in a variety of ways .” And in paragraph [0083]: “ In an implementation, as shown in FIG. 2, the phone feature and the context features may be combined into mixed features, and the mixed features may be provided successively to an attention unit and a decoder to generate acoustic features corresponding to the mixed features. The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .”) Regarding claims 8, 16, 23, and 30, taking claim 16 as exemplary: Zhang and Yuan teach the method, apparatus, non-transitory computer-readable medium, and apparatus of claims 1, 9, 17, and 24 as claimed and specified above. And Zhang shows “in which the at least one processor is further configured to reshape the mixture of channel features to form a … sequence of mixed features.” ( Paragraph [0082]: “At step 1350, a speech waveform corresponding to the text input may be generated based on the phone feature and context features . The speech waveform may be generated in a variety of ways.” And in paragraph [0083]: “In an implementation, as shown in FIG. 2, the phone feature and the context features may be combined into features, and the mixed features mixed may be provided successively to an attention unit and a decoder to generate acoustic features corresponding to the mixed features . The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” In paragraph [0084]: “In another implementation, as shown in FIG. 11, the phone feature and the context features may be combined into first mixed features, and the first mixed features may be provided to an attention unit to obtain first attended mixed features ... The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .” In paragraph [0085]: “In yet another implementation, as shown in FIG. 12, the phone feature and the context features may be combined into first mixed features , and the first mixed features may be provided to the attention unit to obtain first attended mixed features... The acoustic features may then be provided to a vocoder to generate the speech waveform corresponding to the text input .”) And Yuan teaches “one-dimensional” ( Page 562, Paragraph 1: "Figure 3: Illustration of the Locally-enhanced Feed-Forward module. First, patch tokens are projected into a higher dimension . Second, they are restored to “images” in the spatial dimension based on the original positions . Third, a depth-wise convolution is performed on the restored tokens as shown in the yellow region. Then the patch tokens are flattened and projected to the initial dimension. Besides, the class token conducts an identical mapping.") Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure : Wong et al., (US 2023/0118240 A1), part of the prior art made of record, teaches the mixture of features of claims 1, 9, 17, and 24 through the use of “A neural network layer or architecture may comprise a mixture of fixed and learnable parameters” in paragraph [0038] as part of a feed-forward neural network. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHANE D WOOLWINE whose telephone number is (571)272-4138. The examiner can normally be reached M-F 9:30-6:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA HUANG can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHANE D. WOOLWINE Primary Examiner Art Unit 2124 /SHANE D WOOLWINE/Primary Examiner, Art Unit 2124 Application/Control Number: 18/469,909 Page 2 Art Unit: 2124 Application/Control Number: 18/469,909 Page 3 Art Unit: 2124 Application/Control Number: 18/469,909 Page 4 Art Unit: 2124 Application/Control Number: 18/469,909 Page 5 Art Unit: 2124 Application/Control Number: 18/469,909 Page 6 Art Unit: 2124 Application/Control Number: 18/469,909 Page 7 Art Unit: 2124 Application/Control Number: 18/469,909 Page 8 Art Unit: 2124 Application/Control Number: 18/469,909 Page 9 Art Unit: 2124 Application/Control Number: 18/469,909 Page 10 Art Unit: 2124 Application/Control Number: 18/469,909 Page 11 Art Unit: 2124 Application/Control Number: 18/469,909 Page 12 Art Unit: 2124 Application/Control Number: 18/469,909 Page 13 Art Unit: 2124 Application/Control Number: 18/469,909 Page 14 Art Unit: 2124 Application/Control Number: 18/469,909 Page 15 Art Unit: 2124 Application/Control Number: 18/469,909 Page 16 Art Unit: 2124 Application/Control Number: 18/469,909 Page 17 Art Unit: 2124 Application/Control Number: 18/469,909 Page 18 Art Unit: 2124 Application/Control Number: 18/469,909 Page 19 Art Unit: 2124
Read full office action

Prosecution Timeline

Sep 19, 2023
Application Filed
Sep 19, 2025
Examiner Interview Summary
Sep 19, 2025
Applicant Interview (Telephonic)
Apr 07, 2026
Non-Final Rejection mailed — §103, §112
Jun 26, 2026
Response Filed
Sep 30, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12746594
CONTINUOUS CASTING PARAMETER VALUE AND SET-UP CONDITION DETERMINATION USING ARTIFICAL INTELLIGENCE
2y 11m to grant Granted Sep 29, 2026
Patent 12737582
Data Stream-based Computation Unit, Artificial Intelligence Chip, and Accelerator
3y 8m to grant Granted Sep 15, 2026
Patent 12688403
GENERATIVE ARTIFICIAL INTELLIGENCE (AI) SYSTEM
3y 4m to grant Granted Jul 21, 2026
Patent 12682251
FEDERATED LEARNING SYSTEM, FEDERATED LEARNING APPARATUS, FEDERATED LEARNING METHOD, AND FEDERATED LEARNING PROGRAM
2y 11m to grant Granted Jul 14, 2026
Patent 12670455
DYNAMICALLY UPDATED USER INTERFACE FOR THREAT INVESTIGATION
4y 3m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
86%
Grant Probability
99%
With Interview (+20.6%)
2y 10m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 384 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month