DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 9 and 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Dependent claims 9 and 18 recite “wherein the third-NN-leg is insufficient to perform the third task without the information that is sourced from the second-NN-leg”. The limitation recites the term insufficient, however, it is unclear how this ‘insufficiency’ is measured. Since this term is considered subjective, there are no such criteria to determine or measure how insufficient a third NN leg can be to perform a task without information sourced from the second NN leg, as the metes and bounds of this term are unclear. Clarification is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-20 are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Chang et al (NPL: “Classification of Infant Sleep/Wake States: Cross-Attention among Large Scale Pretrained Transformer Networks using Audio, ECG, and IMU Data”, as submitted in IDS dated 01/30/2025- hereinafter Chang).
Referring to Claim 1, Chang teaches a computer-implemented method comprising:
executing a multi-leg neural network (NN) comprising a first-NN-leg and a second-NN-leg (see Chang at Abstract: “We employed a 3-branch (audio/ECG/IMU) large scale transformer based neural network (NN) to demonstrate the potential of such multi-modal data. We pretrained each branch independently with its respective modality, then finetuned the model by fusing the pretrained transformer layers with cross-attention”. Therefore, this 3-branch NN is interpreted as the multi-leg NN), wherein:
the first-NN-leg comprises first-NN-leg layers (see Chang at Fig. 2:
PNG
media_image1.png
234
360
media_image1.png
Greyscale
. It can be seen at Fig. 2 a NN with 3 branches, interpreted as legs, each one having layers);
a first layer of the first-NN-leg layers is at a first depth location in the first NN-leg that corresponds with a first depth location in the second-NN-leg (see Chang at Fig. 2:
PNG
media_image2.png
234
360
media_image2.png
Greyscale
. It can be seen that each layer in each branch are at the same depth location in the NN);
a second layer of the first-NN-leg layers is at a second depth location in the first-NN-leg that corresponds with a second depth location in the second-NN-leg (see Chang at Fig. 2:
PNG
media_image2.png
234
360
media_image2.png
Greyscale
. It can be seen that each layer in each branch are at the same depth location in the NN);
information of the first layer of the first-NN-leg layers is sourced from the first depth location in the second-NN-leg (see Chang at Fig. 2: “Each branch is pretrained individually on unlabeled data. During labeled finetuning, the three branches are combined via cross-attention at the feature level. Their outputs are then concatenated and fed into dense layers to output logits”. Further, see page 2, right column: “Cross-attention based fusion techniques were explored in recent papers such as [37] [38], where the attention layers take concatenated features from different modalities as input, and [39], where a single cross-attention layer is used to share information between branches. Our approach is innovative, first, in that it relies on the pretrained transformer layers from each branch, rather than training a feature sharing mechanism from scratch. Second, we fuse the three transformer networks by alternating self and cross-attention at different layers, which both preserves each branch’s transformer features and incorporates attention from other modalities”); and
information of the second layer of the first-NN-leg layers is sourced from the second depth location in the second-NN-leg (see Chang at Fig. 2: “Each branch is pretrained individually on unlabeled data. During labeled finetuning, the three branches are combined via cross-attention at the feature level. Their outputs are then concatenated and fed into dense layers to output logits”. Further, see page 2, right column: “Cross-attention based fusion techniques were explored in recent papers such as [37] [38], where the attention layers take concatenated features from different modalities as input, and [39], where a single cross-attention layer is used to share information between branches. Our approach is innovative, first, in that it relies on the pretrained transformer layers from each branch, rather than training a feature sharing mechanism from scratch. Second, we fuse the three transformer networks by alternating self and cross-attention at different layers, which both preserves each branch’s transformer features and incorporates attention from other modalities”).
Referring to Claim 2, Chang teaches the computer-implemented method of claim 1, wherein the executing comprises: receiving input comprising a first input type and a second input type to the multi-leg NN; and in response to receiving the input, generating via the multi-leg NN an output for a machine learning main task (see Chang at Fig. 2:
PNG
media_image3.png
234
361
media_image3.png
Greyscale
. It can be seen the inputs of the branches as different types such as audio data, ECG data, and IMU data. Further, see p. 4, left column: “The three branches’ outputs are concatenated and passed through three dense layers with ReLU activations to produce binary logits for sleep and wake”. Therefore, the determination of being asleep or awake is the machine learning main task output).
Referring to Claim 3, Chang teaches the computer-implemented method of claim 1, wherein:
the second-NN-leg comprises second-NN-leg layers; a first layer of the second-NN-leg layers is at the first depth location in the second-NN-leg; a second layer of the second-NN-leg layers is at the second depth location in the second-NN-leg; information of the first layer of the second-NN-leg layers is sourced from the first depth location in the first-NN-leg; and information of the second layer of the second-NN-leg layers is sourced from the second depth location in the first-NN-leg (see Chang at Fig. 2:
PNG
media_image2.png
234
360
media_image2.png
Greyscale
“Each branch is pretrained individually on unlabeled data. During labeled finetuning, the three branches are combined via cross-attention at the feature level. Their outputs are then concatenated and fed into dense layers to output logits”. Further, see page 2, right column: “Cross-attention based fusion techniques were explored in recent papers such as [37] [38], where the attention layers take concatenated features from different modalities as input, and [39], where a single cross-attention layer is used to share information between branches. Our approach is innovative, first, in that it relies on the pretrained transformer layers from each branch, rather than training a feature sharing mechanism from scratch. Second, we fuse the three transformer networks by alternating self and cross-attention at different layers, which both preserves each branch’s transformer features and incorporates attention from other modalities”).
Referring to Claim 4, Chang teaches the computer-implemented method of claim 1, wherein the first layer and the second layer are embedding layers, respectively (see Chang at Fig. 3: “IMU and ECG outputs from the other two branches first get linearly projected and concatenated into a 2-channel embedding. The embedding is reduced to 1-channel by a 1-d convolution layer and passed into MHA as key and value”).
Referring to Claim 5, Chang teaches the computer-implemented method of claim 1, wherein:
the first-NN-leg is operable to, responsive to a first type of input, perform a first task that generates a first instance of a type of predictive output; and the second-NN-leg is operable to, responsive to a second type of input, perform a second task that generates a second instance of the type of predictive output (see Chang at Fig. 2:
PNG
media_image4.png
234
360
media_image4.png
Greyscale
. It can be seen the inputs of the branches as different types such as audio data, ECG data, and IMU data. Further, see p. 4, left column: “The three branches’ outputs are concatenated and passed through three dense layers with ReLU activations to produce binary logits for sleep and wake”. Therefore, the determination of being asleep or awake is the machine learning main task output).
Referring to Claim 6, Chang teaches the computer-implemented method of claim 5, wherein:
at least a portion of the first type of input is different from at least a portion of the second type of input; at least a portion of the first task is different from at least a portion of the second task; and the multi-leg NN generates a final instance of the type of predictive output based at least in part on: the first instance of the type of predictive output; and the second instance of the type of predictive output (see Chang at Fig. 2:
PNG
media_image4.png
234
360
media_image4.png
Greyscale
. It can be seen the inputs of the branches as different types such as audio data, ECG data, and IMU data. Further, see p. 6, left column: “In our variation, we concatenate the three outputs from the feature extractor for each branch as shown in Fig. 2, skipping the transformer layers, and pass through the dense network for fine-tuning. As for late fusion, we leave in the transformer layers, triplicate the FC layers to generate logits for each modality, and average the three logits for evaluation. The proposed cross-attention fusion achieved better performance”).
Referring to Claim 7, Chang teaches the computer-implemented method of claim 5, wherein: the multi-leg NN further comprises a third-NN-leg that, responsive to a third type of input, performs a third task that generates a third instance of the type of predictive output (see Chang at Abstract: “We employed a 3-branch (audio/ECG/IMU) large scale transformer based neural network (NN) to demonstrate the potential of such multi-modal data. We pretrained each branch independently with its respective modality, then finetuned the model by fusing the pretrained transformer layers with cross-attention”. Further, see Fig. 2 which shows the different inputs to each branch, interpreted as legs).
Referring to Claim 8, Chang teaches the computer-implemented method of claim 7, wherein: information of the third-NN-leg is sourced from the second-NN-leg; and the multi-leg NN generates a final instance of the type of predictive output based at least in part on: the first instance of the type of predictive output; the second instance of the type of predictive output; and the third instance of the type of predictive output (see Chang at Fig. 2:
PNG
media_image4.png
234
360
media_image4.png
Greyscale
. It can be seen the inputs of the branches as different types such as audio data, ECG data, and IMU data. Further, see p. 6, left column: “In our variation, we concatenate the three outputs from the feature extractor for each branch as shown in Fig. 2, skipping the transformer layers, and pass through the dense network for fine-tuning. As for late fusion, we leave in the transformer layers, triplicate the FC layers to generate logits for each modality, and average the three logits for evaluation. The proposed cross-attention fusion achieved better performance”).
Referring to Claim 9, Chang teaches the computer-implemented method of claim 8, wherein the third-NN-leg is insufficient to perform the third task without the information that is sourced from the second-NN-leg (see Chang at p.2, right column: “Cross-attention based fusion techniques were explored in recent papers such as [37] [38], where the attention layers take concatenated features from different modalities as input, and [39], where a single cross-attention layer is used to share information between branches. Our approach is innovative, first, in that it relies on the pretrained transformer layers from each branch, rather than training a feature sharing mechanism from scratch. Second, we fuse the three transformer networks by alternating self and cross-attention at different layers, which both preserves each branch’s transformer features and incorporates attention from other modalities”. Further, see pp. 5-6, end of right column of p. 5 to beginning of left column of p. 6: “While all datasets are different and comparison of the accuracies is therefore not theoretically justifiable, such a comparison nevertheless supports the conclusion displayed in the top half of the table, viz., that sleep/wake classification performed using three modalities is more accurate than sleep/wake classification performed using only one or two modalities”. Therefore, since the accuracy is significantly improved using the 3 branches (legs) rather than just using one, this is interpreted as the third (or single branch/leg) being insufficient to perform the task without the cross-attention approach of sharing/sourcing information).
Referring to independent Claims 10 and 19, they are rejected on the same basis as independent claim 1, mutatis mutandis, since both are analogous claims.
Referring to dependent Claims 11 and 20, they are rejected on the same basis as dependent claim 2, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 12, it is rejected on the same basis as dependent claim 3, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 13, it is rejected on the same basis as dependent claim 4, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 14, it is rejected on the same basis as dependent claim 5, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 15, it is rejected on the same basis as dependent claim 6, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 16, it is rejected on the same basis as dependent claim 7, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 17, it is rejected on the same basis as dependent claim 8, mutatis mutandis, since both are analogous claims.
Referring to dependent Claim 18, it is rejected on the same basis as dependent claim 9, mutatis mutandis, since both are analogous claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Lin et al, US 2021/0073644 (this art is pertinent as it discloses a multi-branch neural network includes multiple parallel branches (or layers) that provide outputs, which can be combined using techniques such as summation, concatenation, or other operation used to combine the outputs of the parallel branches);
Wang et al , US 2022/0067274 (this art is pertinent as it discloses attention cross knowledge distillation from a teacher model to a student model, wherein the feature mapping of each layer of the student model is approaching feature mapping of the teacher model, and the student model focuses on intermediate layer features of the teacher model and uses the intermediate layer features to guide the student model).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUIS A SITIRICHE whose telephone number is (571)270-1316. The examiner can normally be reached M-F 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LUIS A SITIRICHE/Primary Examiner, Art Unit 2126