Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 6 is objected to because of the following informalities: “to obtained batched individual losses” is recited in line 4. Should read “to obtain batched individual losses”.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 9 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 9, “image data includes one or more of: audio data, image data, video data, lidar data, radar data, infrared data, ultrasonic data, sensor data, temperature data, pressure data, electrophysiological recording” is recited. The scope of “image data” is unclear because the specification identifies these various types of data as different types of “dimensional data”, going so far as to define audio and temperature data as one dimensional data, separate from the two-dimensional image data (Para. 107). It is unclear whether “image data” is intended to refer specifically to image data or more generally to the input/dimensional data of the training method. Moreover, “image data includes one or more of…. Image data” is generally unclear; how could it not? And what is the significance of the other items on the list given that image data is necessarily included?
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-3, 6-11, and 13-14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Jacob et al. (NPL, “Online Knowledge Distillation for Multi-task Learning”, published 2023, pdf included, also cited in IDS).
Regarding claim 1, Jacob teaches a computer-implemented method for training a multi-task neural network, the multi-task neural network being configured to receive an input and to produce multiple outputs, the multi-task neural network input including dimensional data obtained from a sensor (Abstract, “Multi-task learning (MTL) has found wide application in computer vision tasks. We train a backbone network to learn a shared representation for different tasks such as semantic segmentation, depth- and normal estimation.”), the method comprising the following steps: obtaining a training set, the training set including multiple training pairs of a training input and corresponding multiple training outputs (Abstract, “On the NYUv2 and Cityscapes datasets”); and iterating over the training set (Pg. 3, Col. 2, “The loss weights, λi,i = 1,2,...,Nt, are calculated at each training iteration for each task based on the loss values of single- and multi-task models (sec. 3.3.2)”), including, for each training pair of the training set: evaluating the multi-task neural network for the training input in the training pair, to obtain multiple outputs of the multi-task neural network (Pg. 5, Col. 1, “The multi-task network is trained using a linear combination of task-specific losses”, using a loss indicates evaluation of the multi-task neural network), computing multiple individual losses for the obtained multiple outputs of the multi-task neural network, each individual loss of the multiple individual losses indicating a difference between the output of the multi-task neural network and the corresponding training outputs of the multiple training outputs (Pg. 3, Col. 2, “LiMTL is the task-specific loss for the ith head of the multi-task network”), computing task weights, including computing raw task weights from the multiple individual losses, and applying a normalization function to the raw task weights (Pg. 5, Col. 1, “Let the multi-task model loss at any iteration t be LiMTL (t) and the single task loss LiSTL (t) for the ith task. The task weight for the ith task at iteration t is computed as a temperature-scaled softmax function of the ratio of multi-task to single-task loss”; Eq. 3), and adjusting parameters of the multi-task neural network based on the computed individual losses weighted by the task weights (Pg. 3, Col. 2, “The model is trained in an end-to-end fashion by minimizing the following loss function…”).
Regarding claim 2, Jacob teaches all the elements of claim 1, as stated above, as well as wherein the normalization function is a probability distribution function (Pg. 5, Col. 1, “The task weight for the ith task at iteration t is computed as a temperature-scaled softmax function of the ratio of multi-task to single-task loss”).
Regarding claim 3, Jacob teaches all the elements of claim 1, as stated above, as well as wherein the normalization function includes at least one of: SoftMax, Temperature-Scaled SoftMax, soft-margin SoftMax, Taylor SoftMax (Pg. 5, Col. 1, cited above).
Regarding claim 6, Jacob teaches all the elements of claim 1, as stated above, as well as wherein the iterations over the training set are batched, the multiple individual losses being computed for a batch of training pairs and averaged to obtained batched individual losses, the raw task weights being computed from the batched individual losses (Pg. 3, Col. 2, “at each training iteration”; Pg. 6, Col. 1, “We train all models with the AdamW optimizer [62] and the OneCycleLR scheduler[63].The initial learning rate is set to 10-3 and models are trained for 200 epochs for each dataset”, batching is implicit).
Regarding claim 7, Jacob teaches all the elements of claim 1, as stated above, as well as wherein the training input includes at least image data, and wherein: at least one of the multiple training outputs classify an object in the image data, and/or at least one of the multiple training outputs comprise estimated depth for the image data, and/or at least one of the multiple training outputs comprise an object segmentation for the image data (Abstract, “Multi-task learning (MTL) has found wide application in computer vision tasks. We train a backbone network to learn a shared representation for different tasks such as semantic segmentation, depth- and normal estimation.”; Pg. 4, Col. 1, “As part of a system for visual scene understanding we consider multiple pixel-wise classification and regression tasks.”).
Regarding claim 8, Jacob teaches all the elements of claim 1, as stated above, as well as wherein the training input represents a real-world environment of a mechanical agent operating in the environment, the multi-task neural network being trained to infer a state of the environment and/or of the mechanical agent (Abstract, “On the NYUv2 and Cityscapes datasets”, both these datasets represent the view of an agent operating a real-world environment; Pg. 4, Col. 1, “As part of a system for visual scene understanding we consider multiple pixel-wise classification and regression tasks.”).
Regarding claim 9, Jacob teaches all the elements of claim 7, as stated above, as well as wherein the image data includes one or more of: audio data, image data, video data, lidar data, radar data, infrared data, ultrasonic data, sensor data, temperature data, pressure data, electrophysiological recording (Pg. 5, Col. 2, “The images were hand selected from 435,103 video frames to ensure diverse scene content.”).
Regarding claim 10, Jacob teaches all the elements of claim 1, as stated above, as well as wherein the multi-task neural network includes a shared backbone and multiple task-specific heads, the shared backbone being configured to receive the input, the multiple task-specific heads receiving an output of the shared backbone, each one of the multiple outputs being produced by a corresponding one of the multiple task-specific heads (Pg. 3, Col. 2, “Our framework, shown in Fig. 1, consists of a multi-task network with a shared Vision Transformer (ViT) [25] backbone and separate heads for N tasks.”).
Claim 11 corresponds to claim 1 with the addition of an inference phase (Pg. 5, Col. 2, “It contains 2,975 images for training and 500 for testing, respectively.”), it is similarly rejected.
Claim 13 corresponds to claim 1 and is similarly rejected.
Claim 14 corresponds to claim 1 and is similarly rejected.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over Jacob in view of Lin et al. (NPL, “Dual-Balancing for Multi-Task Learning”, published 2023, pdf attached).
Regarding claim 4, Jacob teaches all the elements of claim 1, as stated above. They do not explicitly disclose wherein the multiple task weights are chosen to scale the corresponding multiple individual losses to a constant.
Lin teaches wherein the multiple task weights are chosen to scale the corresponding multiple individual losses to a constant (Pg. 3, “Logarithmic transformation (Eigen et al., 2014; Girshick, 2015) can be used to achieve the same scale for all losses without the availability of {s⋆ t}T t=1. Specifically, since ∇θ,ψt log ℓt(Dt;θ,ψt) = ∇θ,ψtℓt(Dt;θ,ψt) / ℓt(Dt;θ,ψt), it is equivalent to taking gradient of the scaled task loss ℓt(Dt;θ,ψt) / stop gradient(ℓt(Dt;θ,ψt)), where stop gradient(·) is the stop-gradient operation. Note that ℓt(Dt;θ,ψt) / stop gradient(ℓt(Dt;θ,ψt)) has the same scale (i.e., 1) for all tasks”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jacob to incorporate the teachings of Lin to include wherein the multiple task weights are chosen to scale the corresponding multiple individual losses to a constant. Jacob teaches training a multi-task network using a combination of task-specific losses with task weights that are dynamically determined during training. Lin teaches that differing loss scales among tasks can cause a larger loss to dominate the model update and accordingly teaches scaling the respective task losses to achieve the same scale for all tasks. One of ordinary skill in the art would have recognized that modifying Jacob’s training method to select multiple task weights such that the corresponding individual task losses are scaled to the same constant, as taught by Lin, would have balanced the scales of the different task losses, preventing a task with a larger loss from dominating the training of the multi-task network.
Regarding claim 5, Jacob teaches all the elements of claim 1, as stated above. They do not explicitly disclose wherein a stop gradient operator is applied to the multiple task weights.
Lin teaches wherein a stop gradient operator is applied to the multiple task weights (Pg. 3 cited above, the task loss scaling factor is 1 / stop gradient(ℓt), the stop-gradient operation is applied to the loss from which the task weighing factor is generated.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jacob to incorporate the teachings of Lin to include wherein a stop gradient operator is applied to the multiple task weights. Jacob teaches dynamically determining task weights based on task-specific loss values and using the task weights to train the multi-task network. Lin teaches applying a stop gradient operation to the task loss from which the task weighing factor is generated. One of ordinary skill in the art would have understood that applying this technique to the task weighing of Jacob would have prevented the task-weight calculation from introducing an additional gradient path during training, providing the loss-scale balance as taught by Lin.
Claim(s) 12 is rejected under 35 U.S.C. 103 as being unpatentable over Jacob in view of Kendall et al. (NPL, “Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics”, published 2018, pdf attached, also cited in IDS).
Regarding claim 12, Jacob teaches all the elements of claim 11, as stated above, as well as wherein the input data represents a real-world environment of a mechanical agent, operating in the environment, the multi-task neural network being trained to infer a state of the environment and/or of the mechanical agent and includes an inference phase (See analysis of claims 8 and 11 above).
Jacob does not explicitly disclose that the inference phase includes deriving a control function of the mechanical agent from one or more of the multiple outputs of the multi-task neural network, the control function configured to be executed by the mechanical agent interacting with the real-world environment. However, the datasets for training are from the point of view that an agent would have in a real-world environment.
Kendall teaches wherein the input data represents a real-world environment of a mechanical agent, operating in the environment, the multi-task neural network being trained to infer a state of the environment and/or of the mechanical agent, as well as the importance of multi-task learning for visual scene understanding in robotics (Pg. 1, Cols. 1-2, “Multi task learning of visual scene understanding is of crucial importance in systems where long computation run-time is prohibitive, such as the ones used in robotics. Combining all tasks into a single model reduces computation and allows these systems to run in real-time”).
Kendall does not explicitly disclose deriving a control function of the mechanical agent from one or more of the multiple outputs of the multi-task neural network, the control function configured to be executed by the mechanical agent interacting with the real-world environment. However, they acknowledge the prevalence of using multi-task learning networks in robotics for visual scene understanding (Pg. 1, Col. 1), and they also disclose determining the spatial relationships of pixels (Pg. 2, Col. 2)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jacob to incorporate the teachings of Kendall to include a control function of the mechanical agent from one or more of the multiple outputs of the multi-task neural network, the control function configured to be executed by the mechanical agent interacting with the real-world environment. Jacob discloses a method for multi-task neural network training with a focus on optimizing the training of the network. Kendall is cited within Jacob and teaches that multi-task neural networks are well-known for their usage in visual scene understanding, with robotics being given as an explicit example where multi-task learning can be applied effectively to improve systems that run in real-time (Pg. 1, Cols. 1-2). One of ordinary skill in the art would have understood that implementing the optimized training method of Jacob into the visual scene understanding robotic system disclosed by Kendall would have predictably improved the computer vision performance of the multi-task learning model. Further, a skilled artisan would have found it obvious to derive control commands from perception outputs to drive an agent, as it is a well-established and routine engineering step in robotics.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID A WAMBST whose telephone number is (703)756-1750. The examiner can normally be reached M-F 9-6:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gregory Morse can be reached at (571)272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID ALEXANDER WAMBST/Examiner, Art Unit 2663
/GREGORY A MORSE/Supervisory Patent Examiner, Art Unit 2698