DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-14 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Independent claims 1 and 9 recite, in pertinent part, “a first layer configured to generate feature vectors by projecting training data pairs generated for different tasks to one feature space.” The scope of “training data pairs generated for different tasks” is unclear. In particular, it is unclear whether the recited “training data pairs” are (1) respective training-data pairs associated with respective different tasks, such that each pair pertains to a particular task, or (2) pairs formed from training data corresponding to different tasks, such that a given pair contains training data associated with two or more different tasks. These interpretations impose materially different requirements on the claimed first layer and on the data being projected into the one feature space. The claim does not identify the constituent members of the recited pairs or otherwise specify the relationship between each pair and the “different tasks.”
Accordingly, one of ordinary skill in the art would not be apprised with reasonable certainty of the metes and bounds of the claimed “training data pairs generated for different tasks.”
The specification may describe embodiments in which training samples associated with different tasks are paired with one another; however, limitations appearing in the specification are not imported into the claims where the claim language itself reasonably permits materially different constructions. See MPEP §§ 2173 and 2173.05(e). Applicant may clarify the scope, for example, by expressly reciting that each training-data pair comprises respective training data associated with different tasks, if that is the intended meaning.
Claims 3 and 10 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claims 3 and 10 respectively recite “the feature vector”; however, the claims from which they depend introduce “feature vectors” in the plural and do not previously identify a particular singular feature vector to which “the feature vector” refers. It is therefore unclear whether the recited “the feature vector” refers to one selected feature vector, each respective feature vector associated with a corresponding task, or the plurality of feature vectors collectively. Because these alternatives impose different limitations on the claimed operation, the metes and bounds of the claims cannot be determined with reasonable certainty. See MPEP § 2173.05(e).
Claim 13 is rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claim 13 recites that “all data of each dataset is included in at least one of the training data pairs.” It is unclear whether the limitation requires each individual piece of data in each dataset to be included in at least one respective training data pair, or merely requires the data of each dataset, considered collectively, to be represented in at least one training data pair
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-14 are rejected under 35 U.S.C. 103 as being unpatentable over Sharing Knowledge in Multi-Task Deep Reinforcement Learning, D’Eramo et al; D’Eramo, in view of US Pre-Grant Patent 2021/0232860 (Liu et al; Liu).
Regarding claim 1:
D’Eramo teaches:
1. A multitask learning apparatus comprising:
(D’Eramo, pg. 5, Sect. 3.1)
“The multi-task representation learning problem consists in learning simultaneously a set of T tasks µt, modeled as probability measures over the space of the possible input-output pairs (x,y), with x ∈ X and y ∈ R, being X the input space [i.e. A multitask learning apparatus comprising:].”
2. a first layer configured to generate feature vectors [by projecting training data pairs generated for different tasks to one feature space;]
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Figure 1: (a) The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks [i.e. a first layer configured to generate feature vectors].”
3. a second layer configured to extract a common feature from the projected feature vectors;
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Figure 1: (a) The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks… Note that each block can be composed of arbitrarily many layers. [i.e. a second layer configured to extract a common feature from the projected feature vectors;].”
4. and a third layer configured to draw each individual inference from the extracted common feature,
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Eventually, the shared representation is specialized in block ft and the output yt of the network is computed [i.e. and a third layer configured to draw each individual inference from the extracted common feature].”
5. wherein the first layer and the third layer are task-specific layers, and the second layer is a layer shared between tasks,
(D’Eramo, pg. 17, Sect. C.2, ¶1)
“The network we use consists of 80 ReLu units for each wt,t ∈ {1,...,T} block, with T = 5 [i.e. [i.e. wherein the first layer]. Then, the shared block h consists of one layer with 80 ReLu units and another one with 80 sigmoid units [i.e. and the second layer is a layer shared between tasks,]. Eventually, each ft has a number of linear units equal to the number of discrete actions a(t) i ,i ∈ {1,...,#A(t)} of task µt which outputs the action-value Qt(s,a(t) i ) =yt(s,a(t) i ) =ft(h(wt(s)),a(t) i ),∀s ∈ S(t) [i.e. and the third layer are task-specific layers,].”
6. and the first layer, the second layer, and the third layer perform forward propagation in one artificial neural network.
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Figure 1: (a) The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks. Eventually, the shared representation is specialized in block ft and the output yt of the network is computed.”
PNG
media_image1.png
185
228
media_image1.png
Greyscale
Examiner notes that Figure 1a consists of a single forward propagation flow. Further, the Introduction to D’Eramo states that “…allow us to learn multiple tasks with a single regressor extracting a common representation.”
D’Eramo does not explicitly teach:
1. [a first layer configured to generate feature vectors] by projecting training data pairs generated for different tasks to one feature space;
Liu teaches:
1. [a first layer configured to generate feature vectors] by projecting training data pairs generated for different tasks to one feature space;
(Liu, ¶0060)
“More specifically, PET/MRI pairing information may be used to map the same PET/MRI image pair to the same location in a learned feature space. Then, a subsequent MM mapping in the feature space may be used to retrieve similar cases associated with closely related PET images, which have features that are similar to the MRI image features [i.e. by projecting training data pairs generated for different tasks to one feature space;].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo such that training data corresponding to the different tasks are paired and projected by the respective task-specific mappings into a common feature space, as taught by Liu. Liu recognizes that data originating from different modalities may not be directly comparable and teaches using pairing information and respective mappings to represent such heterogeneous inputs within a shared feature space. Applying this technique to D’Eramo’s task-specific input mappings would predictably permit corresponding training data from the different tasks to be jointly represented and exploited in learning the shared representation, thereby facilitating learning relationships across heterogeneous inputs. Liu expressly identifies the resulting benefit as “facilitation of machine learning for various applications (Liu, ¶0019).”
Regarding claim 2:
D’Eramo and Liu teach:
1. wherein the first layer includes projection encoders.
(Liu, ¶0030)
“Audio training may be achieved by converting audio inputs (e.g., to 2-D Mel-Spectrogram), so that images may be handled in a similar manner. Interference between an image channel and an audio channel may be avoided by employing different weights in the audio encoder and the image encoder. Similarly, different weights may be employed at the audio decoder and the image decoder [i.e. wherein the first layer includes projection encoders.]
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 3:
D’Eramo and Liu teach:
1. wherein the first layer uses individual weight matrices separately allocated to the tasks to generate the feature vector.
(Liu, ¶0062)
“For example, but not by way of limitation, two independent networks may be used to learn mapping weights for the MM and PET images, respectively. This example approach may weight learning interferences between MRI images and PET images.”
Examiner notes that a neural network weight is conventionally understood to be held in a matrix. See patent reference WO 2015011688: “Each of the connections between nodes in successive pairs of layers has an associated weight held in a matrix, and the number of layers and nodes is typically selected or adjusted according to the application the neural network 10 is intended to perform.”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 4:
D’Eramo and Liu teach:
1. wherein the second layer includes a fusion encoder.
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Figure 1: (a) The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks [i.e. wherein the second layer includes a fusion encoder].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 5:
D’Eramo and Liu teach:
1. wherein the second layer uses one weight matrix to extract the common feature.
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Figure 1: (a) The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks [i.e. wherein the second layer uses one weight matrix to extract the common feature].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 6:
D’Eramo and Liu teach:
1. wherein the third layer includes independent classifiers.
(D’Eramo, pg. 6, Sect. 4, ¶2)
“In both cases, the targets are learned for each task µt in its respective output block ft. [i.e. wherein the third layer includes independent classifiers].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 7:
D’Eramo and Liu teach:
1. wherein the training data pairs are generated using a data augmentation technique.
(Liu, ¶0057)
“Watson Text to Speech was used to generate corresponding audio segments for different objects by varying voice model parameters such as expression etc. The new dataset had 72 images and 50 audio segments for each object. In this dataset, 24 images and 10 audio segments from each object category were randomly sampled and used as test data. The remaining images and audio segments were used as training data. Pairing of these images and audio segments was based on 100 underlying object states of the signal generation machine [i.e. wherein the training data pairs are generated using a data augmentation technique].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 8:
D’Eramo and Liu teach:
1. wherein task-specific inference errors and pairwise representation losses are used as loss functions for backpropagation of the artificial neural network.
(D’Eramo, pg. 3, Sect. 3.1, ¶1)
“We assume that the loss function : R × R → [0,1] is 1-Lipschitz in the first argument for every value of the second argument. While this assumption may seem restrictive, the result obtained can be easily scaled to the general case. To use the principal result of this section, for a generic loss function , it is possible to use (·) = (·)/max , where max is the maximum value [i.e. wherein task-specific inference errors… are used as loss functions for backpropagation of the artificial neural network].”
(Liu, ¶0041)
“Their representations are indicated in the shared representation space at 151 and 153. For those media data as inputs, a contrastive loss function is used to simulate the neuron wiring process and the long-term depression process [i.e. and pairwise representation losses].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 9:
D’Eramo and Liu teach:
1. A multitask learning method performed in an artificial neural network including a first layer, a second layer, and a third layer, the multitask learning method comprising:
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks. Eventually, the shared representation is specialized in block ft and the output yt of the network is computed. [i.e. A multitask learning method performed in an artificial neural network including a first layer, a second layer, and a third layer, the multitask learning method comprising:]”
2. generating, by the first layer, feature vectors by projecting training data pairs generated for different tasks to one feature space;
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt.[i.e. generating, by the first layer, feature vectors]
(Liu, ¶0060)
“More specifically, PET/MRI pairing information may be used to map the same PET/MRI image pair to the same location in a learned feature space. Then, a subsequent MM mapping in the feature space may be used to retrieve similar cases associated with closely related PET images, which have features that are similar to the MRI image features [i.e. by projecting training data pairs generated for different tasks to one feature space;].”
3. extracting, by the second layer, a common feature from the projected feature vectors;
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“…to a shared set of layers h which extracts a common representation of the tasks [i.e. extracting, by the second layer, a common feature from the projected feature vectors;]”
4. and drawing, by the third layer, each individual inference from the extracted common feature.
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Eventually, the shared representation is specialized in block ft and the output yt of the network is computed.”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 10:
D’Eramo and Liu teach:
1. wherein the first layer generates the feature vector using individual weight matrices separately allocated to the tasks.
(Liu, ¶0062)
“For example, but not by way of limitation, two independent networks may be used to learn mapping weights for the MM and PET images, respectively. This example approach may weight learning interferences between MRI images and PET images [i.e. wherein the first layer generates the feature vector using individual weight matrices separately allocated to the tasks].”
Examiner notes that a neural network weight is conventionally understood to be held in a matrix. See patent reference WO 2015011688: “Each of the connections between nodes in successive pairs of layers has an associated weight held in a matrix, and the number of layers and nodes is typically selected or adjusted according to the application the neural network 10 is intended to perform.”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 11:
D’Eramo and Liu teach:
(D’Eramo, pg. 6, Sect. 4, ¶1, Figure 1(a))
“Figure 1: (a) The architecture of the neural network we propose to learn T tasks simultaneously. The wt block maps each input xt from task µt to a shared set of layers h which extracts a common representation of the tasks [i.e. wherein the second layer uses one weight matrix to extract the common feature].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 12:
D’Eramo and Liu teach:
1. wherein the training data pairs are generated using a data augmentation technique.
(Liu, ¶0057)
“Watson Text to Speech was used to generate corresponding audio segments for different objects by varying voice model parameters such as expression etc. The new dataset had 72 images and 50 audio segments for each object. In this dataset, 24 images and 10 audio segments from each object category were randomly sampled and used as test data. The remaining images and audio segments were used as training data. Pairing of these images and audio segments was based on 100 underlying object states of the signal generation machine [i.e. wherein the training data pairs are generated using a data augmentation technique].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 13:
D’Eramo and Liu teach:
1. wherein, according to the data augmentation technique, pieces of data randomly extracted from task datasets are paired, and all data of each dataset is included in at least one of the training data pairs.
Liu teaches:
(Liu, ¶0041)
“To simulate the long-term depression process that allows cells to weaken, and eventually eliminate the port connections, the model may allocate a memory that can randomly sample past media data [i.e. wherein, according to the data augmentation technique, pieces of data randomly extracted from task datasets are paired,]”
(Liu, ¶0057)
“The remaining images and audio segments were used as training data. Pairing of these images and audio segments was based on 100 underlying object states of the signal generation machine [i.e. and all data of each dataset is included in at least one of the training data pairs.].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Regarding claim 14:
D’Eramo and Liu teach:
1. wherein task-specific inference errors and pairwise representation losses are used as loss functions for backpropagation of the artificial neural network.
(D’Eramo, pg. 3, Sect. 3.1, ¶1)
“We assume that the loss function : R × R → [0,1] is 1-Lipschitz in the first argument for every value of the second argument. While this assumption may seem restrictive, the result obtained can be easily scaled to the general case. To use the principal result of this section, for a generic loss function , it is possible to use (·) = (·)/max , where max is the maximum value [i.e. wherein task-specific inference errors… are used as loss functions for backpropagation of the artificial neural network].”
(Liu, ¶0041)
“Their representations are indicated in the shared representation space at 151 and 153. For those media data as inputs, a contrastive loss function is used to simulate the neuron wiring process and the long-term depression process [i.e. and pairwise representation losses].”
One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify D’Eramo with Liu. The motivation is the same as claim 1.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL JUSTIN BREENE whose telephone number is (571)272-6320. Examiner
interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-
based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO
Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on 303-297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786 9199 (IN USA OR CANADA) or 571-272-1000.
/P.J.B./ Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129