CTNF 18/833,505 CTNF 87845 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s) feature extraction via dimensional compression (convolution), attention weight generation via cross-modal dimensional compression, applying attention weights (element-wise multiplication), and classification (score vector comparison). This judicial exception is not integrated into a practical application because the cross-attention architecture is a mathematical technique, not a technological improvement. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because they are generic computer components and well-understood data gathering in the art. Step 1: Claims 1-6 are directed to a machine; claims 7-12 to a process; claims 13-18 to a manufacture. Each falls within a statutory category. Step 2A, Prong 1: Taking claim 1 as representative, the claim recites computing intermediate feature values from first and second data via dimensional compression, computing cross-modal attention data (first attention from the second intermediate feature, second attention from the first intermediate feature), applying attention data to compute final feature values, and performing classification using the feature values. These limitations, under their broadest reasonable interpretation, are mathematical calculations – specifically, the computation of attention weights via dimensional compression (convolution), element-wise multiplication of attention with intermediate features, and score-vector-based classification. These fall within the "mathematical concepts" grouping of abstract ideas. MPEP § 2106.04(a)(2)(I). Step 2A, Prong 2: The additional elements “at least one memory” and "at least one processor" are generic computer components that amount to no more than mere instructions to apply the judicial exception on a computer. See MPEP § 2106.05(f). The data acquisition step is insignificant extra-solution activity (mere data gathering). MPEP § 2106.05(g). The specification confirms the computer is any general-purpose computer (¶¶ 0032-0035). The claims do not recite any improvement to computer functionality, a particular machine, a particular transformation, or any other meaningful limitation beyond the abstract idea. Step 2B: The generic computer components are well-understood, routine, and conventional as evidenced by the specification's description of the computer hardware (¶¶ 0032-0035). The data acquisition is well-understood data gathering in the field (¶¶ 0045-0050). No additional element or combination provides an inventive concept. Dependent claims 2-5, 8-11, and 14-17 add further mathematical detail (dimensional compression specifications, link data generation, normalization, weighted assignment) that do not change the § 101 analysis. Claims 6, 12 and 18 specify image features, skeleton features, and motion classification. These are field-of-use limitations that do not integrate the abstract idea into a practical application. MPEP § 2106.05(h). Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-15-aia AIA Claim(s) 1, 2, 7, 8, 13, and 14 is/are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Chen et al., CN 112597884 A (hereinafter "CHEN") . Claims 1, 7, and 13. CHEN discloses a classification apparatus comprising: at least one memory that is configured to store instructions; and at least one processor that is configured to execute the instructions (CHEN ¶¶16-17) to: acquire, for a classification target, first data being a first type of feature value and second data being a second type of feature value (CHEN: "the first modality's multi-dimensional feature data X and the second modality's multi-dimensional feature data Y" (¶35); first data = first modality multi-dimensional feature data X, second data = second modality multi-dimensional feature data Y) ; compute a first intermediate feature value from the first data, and then further computing a first feature value by using the first intermediate feature value (CHEN: "Compress the multi-dimensional feature data of different modalities, and obtain the channel descriptors of the corresponding modalities" (¶36); CHEN: "channel Descriptor ε … has a global receptive field" ¶37. First intermediate = channel descriptor ε x , obtained by per-channel compression of X) ; compute a second intermediate feature value from the second data, and then further computing a second feature value by using the second intermediate feature value (CHEN: "Compress the multi-dimensional feature data of different modalities, and obtain the channel descriptors of the corresponding modalities" (¶36); second intermediate feature value = channel descriptor ε y , obtained by the same per-channel compression of Y) ; perform classification regarding the classification target by using the first feature value, the second feature value, or both thereof (CHEN: "a prediction module for gesture recognition based on the fused feature data" (¶90; module 505 in CHEN Fig. 5 per ¶91); CHEN's "prediction module" under broadest reasonable interpretation because the present specification defines classification as "a process of identifying a class regarding a classification target" where the target may be "any object" including "a person" (SPEC ¶12), and discloses image feature + skeleton feature pairs extracted from image data as the canonical input pair (SPEC ¶16) i.e., the same video + bone modality pairing CHEN uses) ; compute first attention data by using the second intermediate feature value and computing second attention data by using the first intermediate feature value (CHEN: "Calculate the cross-attention value of the corresponding modal based on the channel descriptors of different modalities" (¶40); CHEN: “For the channel descriptor of each modal … β x = softmax( λ 2Relu( λ1 ε x))" (¶¶41-42, formula (2)); “Similarly, through the above steps S110 and S120, the channel descriptor ε y and the cross-attention value β y of the multi-dimensional feature data Y of the second modality can be obtained” (¶45). first attention data = β y computed from ε y per formula (2), second attention data = β x computed from ε x per formula (2)) , wherein the first feature value is computed by using the first intermediate feature value and the first attention data (CHEN: "using the cross-attention value of the multi-dimensional feature data of each mode to strengthen the multi-dimensional feature data of other modes" (¶44); CHEN: X = F( β y, X) = β y • X (146, formula (3)); the first modality multiplied by the attention data β y derived from the second modality) , and wherein the second feature value is computed by using the second intermediate feature value and the second attention data (CHEN: "Y = F( β x, Y) = β x • Y" (¶47, formula (4)); CHEN: "use the cross-attention value β x of the multi-dimensional feature data X of the first mode to strengthen the multi-dimensional feature data Y of the second mode ... use the cross-attention value β Y of the multi-dimensional feature data Y of the second mode to strengthen the multi-dimensional feature data X of the first mode" (¶ 45); the second modality multiplied by the attention data β x derived from the first modality) . PNG media_image1.png 297 540 media_image1.png Greyscale CHEN Fig. 1 - cross-attention mechanism (CHEN ¶35). The first-mode feature data X is reduced to channel descriptor ε x , transformed to cross-attention value β x , and applied to the second-mode feature data Y to produce strengthened Y ^ (and symmetrically X ^ from β y • X). Claim 7 recites the method counterpart of claim 1 (acquisition step, first and second feature extraction steps, classification step, and attention data generation step), and claim 13 recites the non-transitory computer-readable-medium. Claims 7 and 13 are rejected based the same teachings of CHEN, mutatis mutandis . Claim 2, 8 and 14 The classification apparatus according to wherein the computation of the first attention data includes performing dimensional compression on the second intermediate feature value to generate the first attention data having the same number of dimensions as the number of dimensions of the first intermediate feature value, and wherein the computation of the second attention data includes performing dimensional compression on the first intermediate feature value to generate the second attention data having the same number of dimensions as the number of dimensions of the second intermediate feature value (CHEN: "Compress the multi-dimensional feature data of different modalities, and obtain the channel descriptors of the corresponding modalities." (¶36).) (CHEN: "channel Descriptor ε ... has a global receptive field" (¶37). Chen applies the same channel descriptor compression to each modality (¶¶36-37), so ε x and ε y have the same length and shape. Formula 2 is shape preserving, so β y has the same dimensions as ε x and β x has the same dimensions as ε y . ) . Claims 8 and 14 recite the method and non-transitory computer-readable-medium counterparts of claim 2 and are rejected on the same teachings of CHEN. Claims 1, 7, and 13 are alternatively rejected under 35 U.S.C. 102(a)(1) as anticipated by Du et al., CN 114398961 A, (hereinafter "Du” ). Claims 1, 7, and 13. Du discloses a classification apparatus comprising: at least one memory that is configured to store instructions; and at least one processor that is configured to execute the instructions to: acquire, for a classification target, first data being a first type of feature value and second data being a second type of feature value (Du: "Use convolutional neural network and longterm memory network to acquire image and text modal data features respectively" (¶8); first data = image features extracted via Faster-RCNN, second data= text features extracted via GloVe + LSTM) ; compute a first intermediate feature value from the first data, and then further computing a first feature value by using the first intermediate feature value (Du: "The SA(Text) unit and the SA(lmage) unit are processed in parallel to realize the self-attention feature modeling inside the text and the image respectively" (¶23); first intermediate feature value = self-attended image features = output of the SA(lmage) unit) ; compute a second intermediate feature value from the second data, and then further computing a second feature value by using the second intermediate feature value (Du: "The SA(Text) unit and the SA(lmage) unit are processed in parallel to realize the self-attention feature modeling inside the text and the image respectively" (¶23); second intermediate feature value = self-attended text features = output of the SA(Text) unit) ; perform classification regarding the classification target by using the first feature value, the second feature value, or both thereof (Du: "the fusion features are passed to the classifier and combined with the answer text data to predict the results" (¶10); Du expressly recites a classifier, so the claimed "classification" is met) ; compute first attention data by using the second intermediate feature value and computing second attention data by using the first intermediate feature value (Du: "use the MA (Image) unit to use the text features to help obtain the feature information of the key regions of the image. At this time, the text features after the second step of collaborative attention processing provide K, V Vector, image features after self-attention processing as the main body of the co-attention unit." (¶25); first attention data = MA(Image) output, in which the text features (second intermediate) serve as the KN vectors that guide attention learning over the image features) , wherein the first feature value is computed by using the first intermediate feature value and the first attention data (Du: "using the MA (Text) unit, the image features processed by self-attention serve as a 'guidance' to provide the K and V vectors required by the MA unit, the text feature after self-attention processing is used as the Q vector required by the MA unit ... completes the first crossmodal feature interaction" (¶24); second attention data = MA(Text) output, in which the image features (first intermediate) serve as the KN vectors that guide attention learning over the text features) , and wherein the second feature value is computed by using the second intermediate feature value and the second attention data (Du: "two modal features mutually serve as a reference for attention weight learning for deeper feature interaction" (¶9); first feature value = image-side output of MA(Image), combining the image intermediate (as Q) with attention weights derived from text; second feature value= text-side output of MA(Text), combining the text intermediate (as Q) with attention weights derived from image. The two MA units run in parallel and each branch 's attention is derived from the opposite modality's intermediate representation, satisfying the claim's asymmetric cross-attention requirement) . Claims 7 and 13 recite the method and non-transitory computer readable medium counterparts of claim 1, respectively, and are alternatively rejected on the same teachings of Du set forth above . Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-21-aia AIA Claim s 6, 12, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over CHEN in view of Yuan et al., JP 2022-117766 A (hereinafter "YUAN" ) . Claims 6, 12, and 18 CHEN and YUAN disclose the classification apparatus according to claim 1 ,wherein the first data is an image feature value extracted from image data obtained by capturing the classification target, wherein the second data is a skeleton feature value extracted from the image data, and wherein a class of the classification target represents a type of a motion of the classification target. CHEN teaches multi-modal feature extraction including image features, skeleton features, and motion class (CHEN: "for video data, 3D convolution is used to extract features ... For bone data, a fully connected layer can be used to extract bone features." (¶¶76-78, Table 1).) but does not teach a type of motion. However, YUAN teaches image feature value extracted from image data and skeleton feature value extracted from the image data and class as a type of motion (YUAN: "inputting the time-series RGB information constituting the input video to the trained model and calculating a feature vector" (¶32). YUAN: "extracting skeleton information from the input video in time series ... two-dimensional coordinates of parts such as the nose, neck, both shoulders, both elbows, both wrists, both hips, both knees, and both ankles" (¶38). YUAN: "inputting the time-series skeleton information ... to a trained model and calculating a feature vector" (¶39). YUAN: "calculate a score for each action class" (¶48).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to substitute the video/bone modalities of CHEN's cross-attention model with YUAN's RGB image feature and image derived skeleton feature pair to leverage complementary appearance (RGB) and pose (skeleton) cues for improved motion-class recognition accuracy on the same image-captured subject. Claims 12 and 18 recite the method and non-transitory computer-readable-medium counterparts of claim 6, respectively, and are rejected on the same teachings of CHEN and YUAN set out above Allowable Subject Matter 07-43 Claims 3-5, 9-11, and 15-17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, and overcoming the § 101 rejection. Conclusion The prior art made of record but not relied, yet considered pertinent to the applicant’s disclosure, is listed on the PTO-892 form. CN 112101262 A (Ye et al.) discloses multi-feature fused sign language recognition combining RGB color image features (LBP/HOG/RGB) with bone-joint projection features, fused via a three-tier attention network; teaches the image-feature + skeleton-feature + motion-class combination as an alternative to YUAN for any § 103 secondary support on claims 6, 12, and 18. CN 111145913 A (Anhui Iflytek Medical) A discloses a multi-attention classification model employing a row/column-pooled cross-attention block (CAI) between two vector streams; cumulative to CHEN on the asymmetric-cross-attention teaching. CN 111985369 A (NW Polytechnical Univ) discloses a course field multi-modal document classifier built on a "cross-modal attention" convolutional network; teaches intramodal attention on each of image and text branches and bilinear-pooling cross-modal fusion. CN 109460707 A (South China Univ. of Technology) discloses multi-modal action recognition combining video, optical-flow, and human-skeleton information into a deep neural network with an attention-based pooling layer. US 2024/0202532 A1 (Salah et al.) discloses an information processing apparatus that derives per-modality weights from feature values of each modality plus object identifying information, then predicts an object attribute from a weighted concatenation of the per-modality feature values. CN 113902995 A (Univ. Sci. & Tech. China) discloses a multi-modal human behavior recognition designed for occlusion robustness, combining RGB image features with complementary modality data via an attention driven recognition pipeline. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ross Varndell whose telephone number is (571)270-1922. The examiner can normally be reached M-F, 9-5 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’Neal Mistry can be reached at (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Ross Varndell/Primary Examiner, Art Unit 2674 Application/Control Number: 18/833,505 Page 2 Art Unit: 2674 Application/Control Number: 18/833,505 Page 3 Art Unit: 2674 Application/Control Number: 18/833,505 Page 4 Art Unit: 2674 Application/Control Number: 18/833,505 Page 5 Art Unit: 2674 Application/Control Number: 18/833,505 Page 6 Art Unit: 2674 Application/Control Number: 18/833,505 Page 7 Art Unit: 2674 Application/Control Number: 18/833,505 Page 8 Art Unit: 2674 Application/Control Number: 18/833,505 Page 9 Art Unit: 2674 Application/Control Number: 18/833,505 Page 10 Art Unit: 2674 Application/Control Number: 18/833,505 Page 11 Art Unit: 2674 Application/Control Number: 18/833,505 Page 12 Art Unit: 2674