DETAILED ACTION
This Office Action is sent in response to the Applicant’s Communication received on 06/06/2026 for application number 18/447,675. The Office hereby acknowledges receipt of the following and placed of record in file: Specification, Drawings, Abstract, Oath/Declaration, IDS, and Claims.
Claims 1-5 are amended.
Claims 1-5 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
35 USC 112
On pages 6 and 7 of the remarks section, the Applicant argues that the present amendment removes that alleged ambiguity. Current claim 4 no longer recites the former language, and instead recites that the first model and the second model are
configured as one integrated model, and that the integrated model is trained to extract common feature information by receiving the common feature vector of the previous step and the input data corresponding to the current task. Accordingly, the basis of 112(b) rejection is no longer present.
In light of the newly amended claim limitations, the Examiner finds the Applicant’s argument persuasive. Therefore, the 35 USC 112 rejection is withdrawn.
35 USC 101
In summary of pages 7 and 8 of the remarks section, the Applicant argues that amended claim 1 now recites the concrete components and operations that provide that
technical improvement. The disclosed framework improves execution of asynchronous tasks by serially integrating past and present information, including propagating a current-step common feature to a next step, and may improve task execution even when some input data is unavailable. Claims 2-5 further narrow that architecture by reciting significant features common to multiple tasks, weighted summing using a hyperparameter or learned weight, an integrated model embodiment, and use of the previous-step common feature vector to enable execution when a specific input is not received. The amended claims do not merely surround an abstract analysis with data gathering and outputting. Rather, they recite the particular machine-learning processing architecture disclosed as improving asynchronous multi-task learning performance.
After further consideration, in light of the newly amended claim limitations, the Examiner finds the Applicant’s argument persuasive. Therefore, the 35 USC 101 rejection is withdrawn.
35 USC 102 and 103
On pages 8 and 9 of the remarks section, the Applicant argues that Jin's "layers of feature extraction network" to the claimed "plurality of tasks" is not a reasonable interpretation in view of Applicant's Specification, which expressly defines a "task" as a task to be solved through machine learning or a task to be executed through machine learning, and explains that multi-task learning means performing learning on a plurality of tasks using one model. Thus, a plurality of network layers in Jin is not reasonably the same as Applicant's claimed plurality of tasks.
The Examiner respectfully disagrees. First, although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Second, Jin does indeed teach tasks that are solved through machine learning. For clarity, in paragraph 0033 to 0035, Jin teaches, “In step S2, the preprocessed image frame is input into the feature extraction network E-ResNet50 for preliminary feature extraction to obtain the pre-extracted features of the current frame”. In paragraph 0036, Jin teaches the machine learning tasks that are employed to achieve a particular function: “the feature extraction network E-ResNetSO includes a cascaded 7 X 7 convolutional layer, a max pooling layer, four residual blocks, and a receptive field enhancement block. The four residual blocks are each composed of a cascaded 1 X 1, a 3 X3, and a 1 X 1 convolutional layer. The receptive field enhancement block is composed of two 1X1, two 3X3, and one 7X7 convolutional layers connected in parallel, so as to map the preprocessed image frame into a high-dimensional feature space”.
On page 9 of the remarks section, the Applicant further argues that Jin does not disclose all limitations of amended claim 1. Importantly, the Office itself stated, in the rejection of former claim 5, that Jin does not teach "wherein the number and type of input data corresponding to the current task vary for each step." That limitation is now part of amended claim 1. Further, amended claim 1 now recites transmitting the common feature vector of the current step to the next step for integration into a common feature vector of the next step, and the Office Action does not identify any such disclosure in Jin. Accordingly, anticipation has not been established.
The Applicant’s argument has been considered but is moot because of the new ground of rejection given to the amended claim 1.
In summary of pages 9-12, the Applicant further argues that claim 2 requires that the second model receive a plurality of input data corresponding to the current task and extract, from the plurality of input data, the second feature vector including a significant feature common to multiple tasks. The current rejection, however, relies on Han for extracting first and second feature vectors and a current-step feature from semantic feature vectors in a semantic image retrieval context. Han's cited disclosure does not teach or suggest a second model that receives a plurality of input data corresponding to a current task and extracts a significant feature common to multiple tasks in Applicant's asynchronous multi-task learning setting. Present claim 3 now recites that the common feature vector of the current step is generated by weighted summing the first feature vector and the second feature vector, and that the weight used for the weighted summing is a hyperparameter or is determined through learning. The rejection of former claim 3, however, is directed to whether Hajimirsadeghi teaches extracting a first feature vector including common feature information over time by inputting the common feature vector of the previous step to a first model, and inputting the input data corresponding to
the current task to a second model to extract a second feature vector including common feature information of the input data. In other words, the existing rejection addresses the former claim language, not the presently claimed weighted-summing limitation. The rejection of former claim 4 was directed to limitations requiring "determining a weight" and "extracting the common feature vector through an inner product calculation," and
relied on Yan for those limitations. Present claim 4 no longer recites that subject matter. Instead, present claim 4 recites that the first model and the second model are configured as one integrated model, and that the integrated model is trained to extract common feature information by receiving the common feature vector of the previous step and the input data corresponding to the current task. The current rejection is directed to former claim 5, which recited that the number and type of input data corresponding to the current task vary for each step, and relied on Li for that limitation. Present claim 5 no longer recites that limitation. Instead, present claim 5 recites that the common feature vector of the previous step is used to enable execution of the current task when at least one specific input data is not received at the current step. The Office's analysis of Li does not address this amended limitation, and Li's cited discussion of task types, task numbers, and task-weight ranges does not teach or suggest using a previous-step common feature vector to enable current-task execution when input data is missing.
Applicant’s arguments with respect to claim(s) 2-5 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Therefore, the 35 USC 103 prior art rejection has been maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1 and 2 are rejected under 35 U.S.C. 103 as being unpatentable over Jin et al. (CN115861893A, see attached translation), hereinafter Jin, in view of Han et al. (CN111782852A, see attached translation), hereinafter Han, Li et al. (CN111176850A, see attached translation), hereinafter Li, HAJIMIRSADEGHI et al. (US 20200076841 A1), hereinafter Hajimirsadeghi, and Liu et al. (End-to-End Multi-Task Learning with Attention, published 2019).
Regarding claim 1, Jin teaches,
An operation method of a multi-task learning model [Para 0006, This invention provides a multi-vehicle tracking method and system based on visible light images to address the technical problem of the inapplicability of existing multi-target tracking algorithms to multi-vehicle tracking problems], the operation method comprising:
obtaining a common feature vector of a previous step [Para 0034, all image frames input to the feature extraction network need to be preprocessed; Para 0039, Specifically, the pre-extracted features F<sub>1,t-1</sub> from the previous frame image are extracted from the feature library G];
receiving input data (Para 0033-0035, all image frames input to the feature extraction network) corresponding to a current task executed in a current step from among a plurality of tasks (Para 0036, layers of feature extraction network) [Para 0033-0035, In step S1, the current input image frame is preprocessed. Specifically, based on TransTrack, for the current image frame I<sub>t</sub> to be tracked, all image frames input to the feature extraction network need to be preprocessed. The preprocessing operations include: size resizing, center cropping, horizontal flipping, and normalization… In step S2, the preprocessed image frame is input into the feature extraction network E-ResNet50 for preliminary feature extraction to obtain the pre-extracted features of the current frame; Para 0036, the feature extraction network E-ResNetSO includes a cascaded 7 X 7 convolutional layer, a max pooling layer, four residual blocks, and a receptive field enhancement block. The four residual blocks are each composed of a cascaded 1 X 1, a 3 X 3, and a 1 X 1 convolutional layer. The receptive field enhancement block is composed of two 1X1, two 3X3, and one 7X7 convolutional layers connected in parallel, so as to map the preprocessed image frame into a high-dimensional feature space];
extracting an output feature vector (Para 0047, vector features) corresponding to the current task based on the common feature vector of the current step [Para 0047, The extracted feature maps are then input into the channel splitting and adjustment module again to achieve a transformation from image space to vector space, resulting in vector features containing local information. These vector features are then fused with the global features to obtain the output of the context-aware coding layer that takes into account both global and local features];
and outputting output data (Para 0054, output the detected target position) corresponding to the current task based on the output feature vector corresponding to the current task [Para 0054, The first decoding layer is mainly used to detect the aircraft target in the current frame and output the detected target position B<sub>Dt</sub> and the corresponding feature F<sub>DL1-4, Dt</sub> in the current frame].
Jin teaches the limitations of claim 1 including the common feature vector of the previous step (Jin, para 0034 and 0039) and the common feature vector of the current step (Jin, Para 0038).
Jin does not teach wherein the number and type of input data corresponding to the current task vary for each step; extracting a first feature vector including common feature information over time by inputting the common feature vector of the previous step to a first model; inputting the input data corresponding to the current task to a second model and extracting a second feature vector including common feature information of the input data corresponding to the current task; extracting feature of current step based on the first feature vector and the second feature vector.
Han teaches,
extracting feature of current step (Para 0022, to obtain similar semantic feature vectors) based on the first feature vector and the second feature vector [Para 0021, (4) Use the final CNN-RNN network model to extract the text features of the query image and extract its corresponding semantic feature vector; Para 0022, (5) Use the cosine similarity comparison method to compare the semantic feature vector of the query image with the semantic feature vector of other images in the image library to obtain similar semantic feature vectors].
Han is analogous to the claimed invention as they both relate to deep learning feature extraction. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Han and provide extracting features based on first feature vector and second feature vector in order to [Han, para 0025] effectively extract high level concepts while gaining the benefit of the various learning tasks.
Jin-Han does not teach wherein the number and type of input data corresponding to the current task vary for each step; extracting a first feature vector including common feature information over time by inputting the common feature vector of the previous step to a first model; inputting the input data corresponding to the current task to a second model and extracting a second feature vector including common feature information of the input data corresponding to the current task;
Li teaches,
wherein the number and type of input data corresponding to current task (Abstract, target operation) vary for each step (Abstract, when a target operation is detected) [Abstract, when a target operation is detected, obtaining a target value corresponding to the target operation; obtaining a target task from a multidimensional array based on the target value, and adding the target task to the data pool; wherein the multidimensional array includes task types, the number of tasks corresponding to different task types, and a range of task weight values corresponding to different task types; the target value is a random integer value within a preset range; the preset range is determined based on the range of task weight values. The technical solution of this invention predetermines each task and adds each task to a data pool, then retrieves tasks based on the established data pool].
Li is analogous to the claimed invention as they both relate to multi-task machine learning. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Li and provide wherein the number and type of input data corresponding to the current task vary for each step in order to [Li, abstract] avoid problems of hot data and slow system response, and improve the user experience.
Jin-Li-Han teach the above limitations of claim 1 including the extracting of the first feature vector (Han, Para 0019) and the extracting of the second feature vector (Han, Para 0021).
Jin-Li-Han do not teach extracting first feature vector including common feature information over time by inputting the common feature vector of the previous step to a first model; inputting the input data corresponding to the current task to a second model and extracting the second feature vector including common feature information of the input data corresponding to the current task.
Hajimirsadeghi teaches,
extracting first feature vector (Para 0207, generates dense feature vector) including common feature information over time (Para 0207, previous sparse feature vectors) by inputting the common feature vector of the previous step (Para 0207, from previous recurrent steps) to a first model (Para 0152, Each recurrent step may contain an MLP); inputting the input data corresponding to the current task (Para 0207, sparse feature vector 1123 as direct input to recurrent step) to a second model (Para 0152, Each recurrent step may contain an MLP) and extracting the second feature vector (Para 0207, dense feature vector) including common feature information of the input data (Para 0207, sparse feature vector) corresponding to the current task (Para 0207, the current log message) [Para 0152, RNN 720 contains multiple recurrent steps, such as 721-723. Each recurrent step may contain an MLP… Each recurrent step 721-723 corresponds to a sequential time step, such as one for each packet of a network flow; Para 0207, In step 1206, the encoder RNN outputs a respective embedded feature vector that is based on features of the current log message and log messages that occurred earlier in the sequence of related log messages. For example, recurrent step 1133 generates dense feature vector 1143 based on sparse feature vector 1123 as direct input to recurrent step 1133 and also based on cross activation by internal state from previous recurrent steps 1131-1132 that is based on previous sparse feature vectors 1121-1122. Thus, feature embedding into dense feature vector 1143 is contextually based on multiple log messages of original log sequence 1110].
Hajimirsadeghi is analogous to the claimed invention as they both relate to deep learning feature extraction. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Hajimirsadeghi and provide extracting feature vectors by inputting information into models in order to [Han, para 0025] effectively extract high level concepts while gaining the benefit of the various learning tasks.
Jin-Li-Han-Hajimirsadeghi teach the above limitations of claim 1 including the common feature vector of the current step (Jin, Para 0038).
Liu teaches
transmitting feature vector (Sect 3.2 para 2,
a
^
i
(
j
-
1
)
) to the next step (Sect 3.2 para 2, subsequent attention modules in block j) for integration into a common feature vector of the next step (Sect 3.2 para 2, formed by a concatenation of the shared features) [Sect 3.1, para 2, Figure 2 shows a detailed visualisation of our network… As shown, each attention module learns a soft attention mask, which itself is dependent on the features in the shared network at the corresponding layer. Therefore, the features in the shared network, and the soft attention masks, can be learned jointly to maximise the generalization of the shared features across multiple tasks, whilst simultaneously maximizing the task-specific performance due to the attention masks; Sect 3.2 para 2, As shown in Figure 2, the first attention module in the encoder takes as input only features in the shared network. But for subsequent attention modules in block j, the input is formed by a concatenation of the shared features u(j), and the task-specific features from the previous layer
a
^
i
(
j
-
1
)
].
Liu is analogous to the claimed invention as they both relate to multi-task learning. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Liu and provide integrating into a feature vector of the next step in order to [Liu, Sect 1, para 6] enable much more expressive combinations of features to be learned for generalisation across tasks, whilst still allowing for discriminative features to be tailored for each individual task.
Regarding claim 2, Jin-Li-Han-Hajimirsadeghi-Liu teach the limitations of claim 1 including the second model (Hajimirsadeghi, Para 0152), the current task (Jin, Para 0036), the plurality of input data (Jin, Para 0033-0035), and the second feature vector (Han, Para 0021).
Liu further teaches,
wherein model receives a plurality of input data corresponding to current task [Sect 3.2 para 2, As shown in Figure 2, the first attention module in the encoder takes as input only features in the shared network] and extracts feature vector including a significant feature common to multiple tasks [Sect 1, para 6, each attention mask automatically determines the importance of the shared features for the respective task; Sect 3.1, para 1, the shared network learns a compact global feature pool across all tasks].
Liu is analogous to the claimed invention as they both relate to multi-task learning. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Liu and provide extracting significant features common to multiple tasks in order to [Liu, Sect 1, para 6] augment learning by allowing learning of both task shared and task-specific features in a self-supervised, end to-end manner.
Claim(s) 3 is rejected under 35 U.S.C. 103 as being unpatentable over Jin in view of Han, Li, Hajimirsadeghi, and Liu, and in further view of Zhi et al. (US 20210407667 A1), hereinafter Zhi.
Regarding claim 3, Jin-Li-Han-Hajimirsadeghi-Liu teach the limitations of claim 2 including the common feature vector of the current step (Claim 1: Han, Para 0022), the first feature vector, and the second feature vector (Claim 1: Han, para 0020-0021).
Zhi teaches,
wherein feature vector is generated by weighted summing first feature vector and second feature vector [Para 0095, As used herein, a classification can refer to a vector of values; Para 0107, the aggregator 780 calculates an element-wise, weighted sum of the feature vectors 752, 762, and 772 to produce the classification 704], and wherein a weight used for the weighted summing is a hyperparameter (alternate) or is determined through learning [Para 0107, The weights may be set… based on a learned training technique].
Zhi is analogous to the claimed invention as they both relate to feature vector manipulation in machine learning systems. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Zhi and provide weighted summing improve decision making systems by merging a diversity objectives.
Claim(s) 4 is rejected under 35 U.S.C. 103 as being unpatentable over Jin in view of Han, Li, Hajimirsadeghi, and Liu, and in further view of Cuevas Juarez et al. (US 20240303969 A1), hereinafter Cuevas.
Regarding claim 4, Jin-Li-Han-Hajimirsadeghi-Liu teach the limitations of claim 2 including the first model, the second model (Claim 1: Hajimirsadeghi, Para 0152), the common feature vector (Claim 1: Jin, Para 0047), and the input data corresponding to the current task (Claim 1: Li, Abstract).
Cuevas teaches,
wherein the first model and the second model are configured as one integrated model [Para 0058, The attribute prediction model 106 is composed of a first part and a second part. The first part is shared feature vector extraction layers for generating (extracting) a shared feature vector from a product image. The second part is composed of a plurality of layers (attribute specific layers for predicting and outputting values indicating a product type and a plurality of attribute values from the generated shared feature vector)], and
wherein the integrated model is trained to extract common feature information (Para 0068, attribute values for each of a plurality of attributes relating to a product) by receiving feature vector (Para 0060, shared feature vector) and input data (Para 0060, input image) [Para 0060, The first part of the attribute prediction model 106 is a learning model for machine learning that applies an image recognition model. As depicted in FIG. 4A, the first part of the attribute prediction model 106 is composed of an input layer L41 and a plurality of layers (a first layer L42 to an N.sup.th layer L45). The first layer L42 to the N.sup.th layer L45 are composed of an intermediate layer, which includes a plurality of convolution layers, and an output layer for classifying and predicting classes, and generate and output a shared feature vector 42 from an input image 41; Para 0068, The second part of the attribute prediction model 106 includes a plurality of layers (a plurality of estimation layers), which are provided in parallel, use the shared feature vector as an input, and output attribute values for each of a plurality of attributes relating to a product, and an output layer that concatenates and outputs the plurality of attribute values outputted from the plurality of layers (layer branches)].
Cuevas is analogous to the claimed invention as they both relate to extracting features in machine learning systems. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Cuevas and provide integrating multiple models to extract features in order to create more comprehensive feature sets when representing aspects of data.
Claim(s) 5 is rejected under 35 U.S.C. 103 as being unpatentable over Jin in view of Han, Li, Hajimirsadeghi, and Liu, and in further view of Milner (US 6993483 B1), hereinafter Milner.
Regarding claim 5, Jin-Li-Han-Hajimirsadeghi-Liu teach the limitations of claim 1 including the common feature vector of the previous step (Hajimirsadeghi, Para 0207) and the current task (Jin, paras 0033-0036).
Milner teaches,
wherein feature vector is used to enable execution of task when at least one specific input data is not received at the current step [Col 3, lines 58-62, a feature vector estimator arranged, in operation, to receive transmitted feature vectors and responsive to said indication from the missing feature vector detector to estimate a replacement feature vector].
Milner is analogous to the claimed invention as they both relate to utilizing features in computing systems. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jin’s teachings to incorporate the teachings of Milner and provide executing a task when data is not received in order to [Milner, Col 5, lines 33-44] improve the performance of a system by ensuring that missing pieces of data are accounted for and handled appropriately.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYED RAYHAN AHMED whose telephone number is (571)270-0286. The examiner can normally be reached Mon-Fri ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYED RAYHAN AHMED/ Examiner, Art Unit 2126
/DAVID YI/ Supervisory Patent Examiner, Art Unit 2126