CTNF 18/963,359 CTNF 80450 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. This judicial exception is not integrated into a practical application because the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are directed to the abstract idea of processing information using mathematical relationships and mental processing concepts, including using one or more neural networks to generate image object features from audio signals, applying alignment bias and temporal bias in multi-head self-attention, and generating video based on the resulting features. Although the claims recite a neural network, an audio encoder, a decoder, and attention mechanisms, these limitations are recited only at a high level of functional result. The additional elements do not integrate the abstract idea into a practical application because they merely use generic computing components as tools to perform the abstract processing. The recited “variable sample rate,” “alignment bias,” “temporal bias,” “cross-modal multi-head self-attention,” “causal multi-head self-attention,” “sparse multi-head self-attention,” and “autoregressive motion decoder” are described only in terms of the results they achieve, not in terms of a specific improvement to the functioning of the computer or another technical field. The claims do not recite a particular machine configuration or a specific technical solution to a technological problem; rather, they amount to instructions to use neural-network-based analysis to generate image object features and video. The additional claim elements, considered individually and as an ordered combination, do not amount to significantly more than the abstract idea itself. The recited processors, memories, and neural networks are generic components performing their ordinary functions. Use of known machine learning techniques, including attention mechanisms and decoders, does not transform the abstract idea into patent-eligible subject matter when claimed at a functional level. Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15-aia AIA Claim s 1, 4, 7, 8, 11 and 14, 15, 18, 20 are rejected under 35 U.S.C. 102 a2 as being anticipated by Kim et al (US 20250117185 A1) (hereinafter Kim) . Regarding claims 1, 8 and 15, Kim discloses a processor, method and system comprising one or more processors, one or more circuits (see figs. 1-2, and paragraph 39) using one or more neural networks (see paragraph 39) to generate image object features based, at least in part, on a variable sample rate of one or more audio signals (i.e., the on-screen classifier module 206 includes a first convolutional neural network that extracts image embedding for the image frames…(see paragraph 78)…. In some embodiments, the second convolutional neural network may include multiple layers and may be trained to analyze audio, e.g., audio spectrograms corresponding to the video frames…(see paragraph 79)… The uncompressed file 504 is modified to obtain a converted file 506 that is suitable as input to the audio separation model. The converted file 506 may include a video at 1080p resolution, a video recording with 4K resolution etc., and 1 frame per second by selecting a single frame of the video every second (or 2, 3 or other number of frames per second, lower than the 25/30/60 fps original video, using frame sampling) (see paragraph 100). Also refer to paragraphs 55, 96 and 106). Regarding claims 4, 11, and 18, Kim discloses a method, a processor and system as disclosed above (see claims 1, 8 and 15 rejections) further to generate a video based, at least in part, on the image object features (see paragraphs 78 and 79). Regarding claims 7, 14 and 20, Kim discloses a method, a processor and system as disclosed above (see claims 1, 8 and 15 rejections) wherein the object is a face (see fig. 7) and wherein the one or more audio signals are speech (see paragraphs 44 and 58) . Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 2, 3, 5, 6, 9, 10, 12, 13, 16, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kim in view of Galvin (US 20250363385 A1) . Regarding claims 2, 9, and 16, Kim discloses a method, a processor and system as disclosed above (see claims 1, 8 and 15 rejections). Kim does not, but Galvin does specifically disclose a method, a processor and system wherein the variable sample rate of one or more audio signals is used to apply an alignment bias learned from encoded audio data of the one or more audio signals as part of cross-modal multi-head self-attention in a decoder of the one or more neural networks (i.e., the processed inputs then enter the fusion layer 1430. The feature alignment unit 1431 uses dynamic time warping for temporal alignment of features from different modalities, ensuring synchronization of time-varying inputs. For spatial alignment, it employs a spatial transformer network to align features in a common coordinate system. The cross-modal attention mechanism 1432 implements a multi-head attention architecture. It computes attention weights between features from different modalities using scaled dot-product attention, allowing the model to focus on relevant cross-modal interactions. The feature concatenation unit 1433 combines the aligned and attended features from different modalities into a single tensor, preserving the temporal and spatial structure of the inputs where applicable. The multimodal transformer 1434 consists of a stack of transformer encoder layers. Each layer includes multi-head self-attention mechanisms and position-wise feed-forward networks. The self-attention allows the model to capture long-range dependencies within and across modalities, while the feed-forward networks introduce non-linearity and increase the model's capacity to learn complex patterns. The dimensionality reduction component 1435 employs a combination of linear projections and non-linear activations to reduce the feature space. It uses a bottleneck architecture similar to autoencoders, where the central layer represents the reduced dimensionality. The normalization component 1436 applies layer normalization to the reduced feature representations. This normalization stabilizes the learning process by normalizing the inputs across the feature dimension, with learnable gain and bias parameters) (see paragraph 226. Also refer to Abstract and paragraph 127). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Kim with the teaching of Galvin to arrive at the claimed invention. A motivation for doing would have been to improve computational efficiency. Regarding claims 3, 10, and 17, Kim discloses a method, a processor and system as disclosed above (see claims 1, 8 and 15 rejections). Kim does not, but Galvin does specifically disclose a method, a processor and system wherein to generate image object features is further based on application of a temporal bias learned from encoded audio data of the one or more audio signals as part of causal multi-head self-attention in a decoder of the one or more neural networks (i.e., the processed inputs then enter the fusion layer… the feature alignment unit 1431 uses dynamic time warping for temporal alignment of features from different modalities, ensuring synchronization of time-varying inputs. For spatial alignment, it employs a spatial transformer network to align features in a common coordinate system. The cross-modal attention mechanism 1432 implements a multi-head attention architecture. It computes attention weights between features from different modalities using scaled dot-product attention, allowing the model to focus on relevant cross-modal interactions. The feature concatenation unit 1433 combines the aligned and attended features from different modalities into a single tensor, preserving the temporal and spatial structure of the inputs where applicable. The multimodal transformer 1434 consists of a stack of transformer encoder layers. Each layer includes multi-head self-attention mechanisms and position-wise feed-forward networks. The self-attention allows the model to capture long-range dependencies within and across modalities, while the feed-forward networks introduce non-linearity and increase the model's capacity to learn complex patterns. The dimensionality reduction component 1435 employs a combination of linear projections and non-linear activations to reduce the feature space. It uses a bottleneck architecture similar to autoencoders, where the central layer represents the reduced dimensionality. The normalization component 1436 applies layer normalization to the reduced feature representations. This normalization stabilizes the learning process by normalizing the inputs across the feature dimension, with learnable gain and bias parameters) (see paragraph 226. Also refer to Abstract and paragraph 127. The abstract discloses Modality agnostic Large Codeword Model which includes causal attention and cross-modal bending). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Kim with the teaching of Galvin to arrive at the claimed invention. A motivation for doing would have been to improve computational efficiency. Regarding claims 5, 12, and 19, Kim discloses a method, a processor and system as disclosed above (see claims 1, 8 and 15 rejections). Kim does not, but Galvin does specifically disclose a method, a processor and system wherein the one or more audio signals are encoded based, at least in part, on application of a sparse multi-head self-attention in an audio encoder in the one or more neural networks (Same as above. Refer to abstract, paragraphs 121, 127 (multi-head attention mechanism) and paragraph 226. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Kim with the teaching of Galvin to arrive at the claimed invention. A motivation for doing would have been to improve computational efficiency. Regarding claims 6 and 13, Kim discloses a processor and method as disclosed above (see claims 1, 8 and 15 rejections). Although Kim discloses a processor and method as disclosed above. Kim does not specifically disclose a method and processor wherein the one or more neural networks include an autoregressive motion decoder to generate the image object features. However, Galvin discloses a method and processor wherein the one or more neural networks include an autoregressive motion decoder to generate the image object features (i.e., Video generator 1740 consists of frame sequence generator 1741 and a motion synthesis subsystem 1742 . frame sequence generator 1741 creates a series of video frames based on the input representation, establishing the visual content of the video output. Motion synthesis subsystem 1742 ensures smooth and realistic motion between frames) (see paragraph 253. Also refer to paragraph 140). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Kim with the teaching of Galvin to arrive at the claimed invention. A motivation for doing would have been to creating coherent and natural-looking video sequences. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PIERRE LOUIS DESIR whose telephone number is (571)272-7799. The examiner can normally be reached Monday-Friday 9AM-5:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659 Application/Control Number: 18/963,359 Page 2 Art Unit: 2659 Application/Control Number: 18/963,359 Page 3 Art Unit: 2659 Application/Control Number: 18/963,359 Page 4 Art Unit: 2659 Application/Control Number: 18/963,359 Page 5 Art Unit: 2659 Application/Control Number: 18/963,359 Page 6 Art Unit: 2659 Application/Control Number: 18/963,359 Page 7 Art Unit: 2659 Application/Control Number: 18/963,359 Page 8 Art Unit: 2659 Application/Control Number: 18/963,359 Page 9 Art Unit: 2659