DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
The present application is being examined under the claims filed 03/26/2024.
Claims 1-21 are pending.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 12/05/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 01/15/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 10/02/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are as follows, and a review of the specification reveals what appears to be the corresponding structure disclosed:
[*Examiner notes: Applicant has provided that (Specification paragraph 165) “The involved module described in the embodiment of the present disclosure can be implemented by software or hardware” and (Specification paragraph 166) “The functions described herein above can be performed, at least in part, by one or more hardware logic components. For example, without limitations, an exemplary type of hardware logic component that can be used includes: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), and so on.” Invoking 35 U.S.C. 112(f) necessarily limits the modules as including at least hardware.]
an object obtaining module (claim 11)
(paragraph 6) At least one embodiment of the present disclosure further provides an object processing apparatus, which comprises: an object obtaining module, configured to obtain a target object to be processed
a first determining module (claim 11)
(paragraph 6) a first determining module, configured to determine an object type of the target object
a second determining module (claim 11)
(paragraph 6) a second determining module, configured to determine a task type corresponding to a target task for processing the target object
an object processing module (claims 11, 16)
(paragraph 6) an object processing module, configured to input the target object
a model generation module (claims 17-20)
(paragraph 217) a model generation module, configured to: obtain a plurality of first sample sets, wherein each of the plurality of first sample sets includes a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types; determine a second sample set according to the plurality of first sample sets; and train a multimodal model according to the second sample set to obtain the target model.
PNG
media_image1.png
309
553
media_image1.png
Greyscale
a storage apparatus (claim 13)
(paragraph 157) a storage apparatus 2008 including, for example, a magnetic tape, a hard disk, etc.
a processing apparatus (claim 13)
(paragraph 156) As shown in FIG. 8, the electronic device 2000 can include a processing apparatus (e.g., central processing unit, graphics processing unit, etc.) 2001, which can execute various appropriate actions and processes according to a program stored on a read-only memory (ROM) 2002 or a program loaded from a storage apparatus 2008 into a random access memory (RAM) 2003. In the RAM 2003, various programs and data necessary for the operations of the electronic device 2000 are also stored. The processing apparatus 2001, the ROM 2002, and the RAM 2003 are connected to each other through a bus 2004. An input/output (I/O) interface 2005 is also connected to the bus 904.
PNG
media_image2.png
370
630
media_image2.png
Greyscale
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 8-11 and 16-20 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Regarding Claim 8
Claim 8 recites the limitation "calculating a comprehensive loss value according to the task loss values and the task weights of the plurality of task processing modules". There is insufficient antecedent basis for this limitation in the claim. In particular, there is insufficient antecedent basis for the term “the task weights of the plurality of task processing modules”. For purposes of examination, the examiner interprets the limitation as though it said "calculating a comprehensive loss value according to the task loss values and [[the]] task weights of the plurality of task processing modules". Examiner suggests amending the claim with this language.
Regarding Claims 11 and 16-20
Claim limitations:
(claim 11) an object obtaining module, configured to obtain a target object to be processed
(claim 11) a first determining module, configured to determine an object type of the target object;
(claim 11) a second determining module, configured to determine a task type corresponding to a target task for processing the target object;
(claim 11) an object processing module, configured to input the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model
(claim 16) wherein the object processing module is further configured to:
(claim 17) further comprising a model generation module, configured to
(claims 18-20) wherein the model generation module is further configured to
invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The specification must explicitly disclose the algorithm for performing the claimed function, and simply reciting the claimed function in the specification will not be a sufficient disclosure for an algorithm which, by definition, must contain a sequence of steps. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. See MPEP 2181 II C.
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
Regarding dependent claims
Claims
9-10 are dependent upon claim 8
16-20 are dependent upon claim 11
18-20 are dependent upon claim 17
19-20 are dependent upon claim 18
20 is dependent upon claim 19
and are therefore similarly rejected for including the deficiencies of claims 8, 11, 17, 18, and 19 respectively.
Claim Rejections - 35 USC § 101 — signals per se
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to signals per se.
Regarding Claim 12
Even when a product has a physical or tangible form, it may not fall within a statutory category. For instance, a transitory signal, while physical and real, does not possess concrete structure that would qualify as a device or part under the definition of a machine, is not a tangible article or commodity under the definition of a manufacture (even though it is man-made and physical in that it exists in the real world and has tangible causes and effects), and is not composed of matter such that it would qualify as a composition of matter. See MPEP 2106.03 I. Claim 12 recites a “computer readable medium” which is not limited to non-transitory media. In fact, the specification explicitly includes transitory forms of storage media in the interpretation (see specification paragraph 159).
Claim Rejections - 35 USC § 101 — abstract idea
Claims 1-20 are rejected under 35 U.S.C. 101 for containing an abstract idea without significantly more.
Regarding Claim 1:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes, the claim is to a process.
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
determining an object type of the target object — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to making a judgement about the type of an object in question.
determining a task type corresponding to a target task for processing the target object — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to making a judgement of the appropriate task to perform on an object in question.
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, the claim does not recite additional elements that integrate the judicial exception into a practical application. The additional elements:
An object processing method, wherein the method comprises: obtaining a target object to be processed — This limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)).
inputting the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model; wherein the target model comprises a feature extraction module, a plurality of object segmentation modules and a plurality of task processing modules, different object segmentation modules correspond to different object types, and different task processing modules correspond to different task types — This limitation is directed to mere instructions to apply a judicial exception. Using an ordinary neural network to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the ordinary is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the abstract idea itself?
No, the claim does not recite additional elements which amount to significantly more than the abstract idea itself. The additional elements as identified in step 2A prong 2:
An object processing method, wherein the method comprises: obtaining a target object to be processed — This limitation is recited at a high level of generality and amounts to mere data gathering of storing and retrieving information in memory, which is well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.), which cannot amount to significantly more than the judicial exception.
inputting the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model; wherein the target model comprises a feature extraction module, a plurality of object segmentation modules and a plurality of task processing modules, different object segmentation modules correspond to different object types, and different task processing modules correspond to different task types — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Regarding Claim 2
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which included an abstract idea (see rejection for claim 1). The claim recites the additional limitations:
Step 2A Prong 2:
wherein the inputting the target object, the object type and the task type into the pre-generated target model to obtain the target result output by the target model comprises: inputting the target object into a target segmentation module to obtain a plurality of first object features after segmentation, wherein the target segmentation module comprises an object segmentation module corresponding to the object type; inputting the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module; inputting the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type — This limitation is directed to mere instructions to apply a judicial exception. Using an ordinary neural network to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the ordinary is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the inputting the target object, the object type and the task type into the pre-generated target model to obtain the target result output by the target model comprises: inputting the target object into a target segmentation module to obtain a plurality of first object features after segmentation, wherein the target segmentation module comprises an object segmentation module corresponding to the object type; inputting the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module; inputting the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 3
Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 2 which included an abstract idea (see rejection for claim 2). The claim recites the additional limitations:
Step 2A Prong 2:
wherein the feature extraction module comprises a shared layer and candidate adaptation layers, each of the candidate adaptation layers corresponds to a different task type; the inputting the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module comprises: performing feature extraction on the plurality of first object features according to the shared layer and a target adaptation layer to obtain the second object feature, wherein the target adaptation layer is a candidate adaptation layer corresponding to the task type — This limitation is directed to mere instructions to apply a judicial exception. Using an ordinary neural network to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the ordinary is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the feature extraction module comprises a shared layer and candidate adaptation layers, each of the candidate adaptation layers corresponds to a different task type; the inputting the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module comprises: performing feature extraction on the plurality of first object features according to the shared layer and a target adaptation layer to obtain the second object feature, wherein the target adaptation layer is a candidate adaptation layer corresponding to the task type — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 4
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which included an abstract idea (see rejection for claim 1). The claim recites the additional limitations:
Step 2A Prong 1:
determining a second sample set according to the plurality of first sample sets — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Step 2A Prong 2:
wherein the target model is generated by: obtaining a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types — This limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)).
training a multimodal model according to the second sample set to obtain the target model — This limitation is directed to mere instructions to apply a judicial exception. Using machine learning training to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the training is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the target model is generated by: obtaining a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types — This limitation is recited at a high level of generality and amounts to mere data gathering of storing and retrieving information in memory, which is well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.), which cannot amount to significantly more than the judicial exception.
training a multimodal model according to the second sample set to obtain the target model — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 5
Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which included an abstract idea (see rejection for claim 4). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the determining a second sample set according to the plurality of first sample sets comprises: taking the plurality of first sample sets as the second sample set; or sampling the plurality of first sample set according to task types to obtain the second sample set — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 6
Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 5 which included an abstract idea (see rejection for claim 5). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the sampling the plurality of first sample set according to task types to obtain the second sample set comprises: determining sampling weights according to the task types; sampling each of the plurality of first sample sets according to the sampling weights to obtain a third sample set corresponding to the each of the plurality of first sample sets; determining the second sample set according to the third sample set — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 7
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 6 which included an abstract idea (see rejection for claim 6). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the determining the second sample set according to the third sample set comprises: taking the third sample set as the second sample set; or taking a third sample set with a same task type as a task processing module of the target model as the second sample set — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 8
Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 7 which included an abstract idea (see rejection for claim 7). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the model training step comprises: obtaining a task loss value of each task processing module of the multimodal model according to the second sample set, wherein the task loss value is used to characterize a difference between a prediction result output by the task processing module and the sample result — This limitation is directed to the abstract idea of a mathematical process, and mathematical calculations in particular (MPEP 2106.04(a)(2) I. C.). The claim describes the mathematical operation of calculating a function in words.
calculating a comprehensive loss value according to the task loss values and the task weights of the plurality of task processing modules — This limitation is directed to the abstract idea of a mathematical process, and mathematical calculations in particular (MPEP 2106.04(a)(2) I. C.). The claim describes the mathematical operation of calculating a function in words.
Step 2A Prong 2:
wherein the training a multimodal model according to the second sample set to obtain the target model comprises: performing a model training step cyclically according to the second sample set, until the trained multimodal model is determined to meet a preset iteration stopping condition, and taking the trained multimodal model as the target model; updating, in a case where the multi-modal model is not determined to meet the preset iteration stopping condition according to the comprehensive loss value, parameters of the multi-modal model to obtain a trained multimodal model, and taking the trained multimodal model as a new multimodal model — This limitation is directed to mere instructions to apply a judicial exception. Using generic machine learning training to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the training is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the training a multimodal model according to the second sample set to obtain the target model comprises: performing a model training step cyclically according to the second sample set, until the trained multimodal model is determined to meet a preset iteration stopping condition, and taking the trained multimodal model as the target model; updating, in a case where the multi-modal model is not determined to meet the preset iteration stopping condition according to the comprehensive loss value, parameters of the multi-modal model to obtain a trained multimodal model, and taking the trained multimodal model as a new multimodal model — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 9
Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 8 which included an abstract idea (see rejection for claim 8). The claim recites the additional limitations:
Step 2A Prong 1:
obtaining the task loss value of each task processing module according to the prediction result and the sample result — This limitation is directed to the abstract idea of a mathematical process, and mathematical calculations in particular (MPEP 2106.04(a)(2) I. C.). The claim describes the mathematical operations of calculating a function.
Step 2A Prong 2:
wherein the obtaining a task loss value of each task processing module of the multimodal model according to the second sample set comprises: inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation, wherein the sample segmentation module comprises an object segmentation module corresponding to an object type of the sample object; inputting the plurality of the first sample object features into the feature extraction module to obtain a second sample object feature output by the feature extraction module; inputting the second sample object feature into a sample processing module to obtain a prediction result output by the sample processing module, wherein the sample processing module comprises a task processing module corresponding to the task type —This limitation is directed to mere instructions to apply a judicial exception. Using an ordinary machine learning model to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the machine learning is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the obtaining a task loss value of each task processing module of the multimodal model according to the second sample set comprises: inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation, wherein the sample segmentation module comprises an object segmentation module corresponding to an object type of the sample object; inputting the plurality of the first sample object features into the feature extraction module to obtain a second sample object feature output by the feature extraction module; inputting the second sample object feature into a sample processing module to obtain a prediction result output by the sample processing module, wherein the sample processing module comprises a task processing module corresponding to the task type — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 10
Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 9 which included an abstract idea (see rejection for claim 9). The claim recites the additional limitations:
Step 2A Prong 2:
wherein the inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation comprises: segmenting, in a case where the object type of the sample object is image or audio, to obtain a plurality of sub-sample objects, wherein an overlapping region exists between sub-sample objects that are adjacent to each other; determining the plurality of first sample object features according to the plurality of sub-sample objects — This limitation is directed to mere instructions to apply a judicial exception. Using ordinary transformer operations (e.g. tokenization) to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the transformer operations is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation comprises: segmenting, in a case where the object type of the sample object is image or audio, to obtain a plurality of sub-sample objects, wherein an overlapping region exists between sub-sample objects that are adjacent to each other; determining the plurality of first sample object features according to the plurality of sub-sample objects — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 11:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes, the claim is to a machine.
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
a first determining module, configured to determine an object type of the target object — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to making a judgement about the type of an object in question.
a second determining module, configured to determine a task type corresponding to a target task for processing the target object — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to making a judgement of the appropriate task to perform on an object in question.
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, the claim does not recite additional elements that integrate the judicial exception into a practical application. The additional elements:
An object processing apparatus, comprising: an object obtaining module, configured to obtain a target object to be processed — This limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)).
an object processing module, configured to input the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model; wherein the target model comprises a feature extraction module, a plurality of object segmentation modules and a plurality of task processing modules, different object segmentation modules correspond to different object types, and different task processing modules correspond to different task types — This limitation is directed to mere instructions to apply a judicial exception. Using an ordinary neural network to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the ordinary is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the abstract idea itself?
No, the claim does not recite additional elements which amount to significantly more than the abstract idea itself. The additional elements as identified in step 2A prong 2:
An object processing apparatus, comprising: an object obtaining module, configured to obtain a target object to be processed — This limitation is recited at a high level of generality and amounts to mere data gathering of storing and retrieving information in memory, which is well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.), which cannot amount to significantly more than the judicial exception.
an object processing module, configured to input the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model; wherein the target model comprises a feature extraction module, a plurality of object segmentation modules and a plurality of task processing modules, different object segmentation modules correspond to different object types, and different task processing modules correspond to different task types — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Regarding Claim 12
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which included an abstract idea (see rejection for claim 1). The claim recites the additional limitations:
Step 2A Prong 2:
A computer-readable medium, storing a computer program thereon, wherein the computer program, when executed by a processing apparatus, realizes steps of the method according to claim 1 — This limitation is directed to merely applying an abstract idea using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.04(d)).
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
A computer-readable medium, storing a computer program thereon, wherein the computer program, when executed by a processing apparatus, realizes steps of the method according to claim 1 — Using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 13
Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which included an abstract idea (see rejection for claim 1). The claim recites the additional limitations:
Step 2A Prong 2:
An electronic device, comprising: a storage apparatus, storing a computer program thereon; a processing apparatus, configured to execute the computer program on the storage apparatus to realize steps of the method according to claim 1 — This limitation is directed to merely applying an abstract idea using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.04(d)).
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
An electronic device, comprising: a storage apparatus, storing a computer program thereon; a processing apparatus, configured to execute the computer program on the storage apparatus to realize steps of the method according to claim 1 — Using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 14
Claim 14 is identical to claim 4 except that claim 14 is dependent upon claim 2 instead of claim 1. Accordingly, the same rejection and rationale applies.
Regarding Claim 15
Claim 15 is identical to claim 4 except that claim 15 is dependent upon claim 3 instead of claim 1. Accordingly, the same rejection and rationale applies.
Regarding Claim 16
Claim 16 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 11 which included an abstract idea (see rejection for claim 11). The claim recites the additional limitations:
Step 2A Prong 2:
wherein the object processing module is further configured to: input the target object into a target segmentation module to obtain a plurality of first object features after segmentation, wherein the target segmentation module comprises an object segmentation module corresponding to the object type; input the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module; input the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type — This limitation is directed to mere instructions to apply a judicial exception. Using an ordinary neural network to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the ordinary is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
wherein the object processing module is further configured to: input the target object into a target segmentation module to obtain a plurality of first object features after segmentation, wherein the target segmentation module comprises an object segmentation module corresponding to the object type; input the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module; input the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 17
Claim 17 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 11 which included an abstract idea (see rejection for claim 11). The claim recites the additional limitations:
Step 2A Prong 1:
determine a second sample set according to the plurality of first sample sets — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Step 2A Prong 2:
further comprising a model generation module, configured to: obtain a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types — This limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)).
train a multimodal model according to the second sample set to obtain the target model — This limitation is directed to mere instructions to apply a judicial exception. Using machine learning training to apply a judicial exception (see MPEP 2106.05(f)) is insufficient to integrate the judicial exception into a practical application. Even if the training is implemented on a generic computer (see MPEP 2106.05(f)(2), 2106.04(d)), the limitation does not integrate the judicial exception into a practical application.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2.
Step 2B:
The additional elements as identified in step 2A prong 2:
further comprising a model generation module, configured to: obtain a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types — This limitation is recited at a high level of generality and amounts to mere data gathering of storing and retrieving information in memory, which is well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.), which cannot amount to significantly more than the judicial exception.
train a multimodal model according to the second sample set to obtain the target model — Mere instructions to apply a judicial exception (see MPEP 2106.05(f)) and using a generic computer as a tool (see MPEP 2106.05(f)(2), 2106.05(d)) cannot amount to significantly more than the judicial exception itself.
Thus, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 18
Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 17 which included an abstract idea (see rejection for claim 17). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the model generation module is further configured to: take the plurality of first sample sets as the second sample set; or sample the plurality of first sample set according to task types to obtain the second sample set — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 19
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 18 which included an abstract idea (see rejection for claim 18). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the model generation module is further configured to: determine sampling weights according to the task types; sample each of the plurality of first sample sets according to the sampling weights to obtain a third sample set corresponding to the each of the plurality of first sample sets; determine the second sample set according to the third sample set— This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 20
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 19 which included an abstract idea (see rejection for claim 19). The claim recites the additional limitations:
Step 2A Prong 1:
wherein the model generation module is further configured to: take the third sample set as the second sample set; or take a third sample set with a same task type as a task processing module of the target model as the second sample set — This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). The limitation is directed to a mental process because it amounts to evaluating known data to make a selection of data that should go in a set.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1 and 11 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Akbari et al “VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text” herein referred to as Akbari.
Regarding Claim 1
Akbari teaches:
An object processing method, wherein the method comprises: obtaining a target object to be processed;
(page 1 abstract) “Specifically, our Video Audio-Text Transformer (VATT) takes raw signals as inputs[*Examiner notes: target object to be processed] and extracts multi modal representations that are rich enough to benefit a variety of downstream tasks.”
determining an object type of the target object;
(page 4 section 3.1) “VATT operates on raw signals. The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers.”; [*Examiner notes: The object type is determined when it is input to the modality-specific tokenization layer]
determining a task type corresponding to a target task for processing the target object;
(page 6 section 4.1 “Downstream”) “We evaluate the pre-trained Transformers on a variety of downstream tasks: image classification, video action recognition, audio event classification, and zero-shot text-to-video retrieval.”
inputting the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model;
(page 4 section 3.2) “For simplicity, we adopt the most established Transformer architecture[*Examiner notes: pre-generated model] [23], which has been widely used in NLP. Similar to ViT [25], we do not tweak the architecture so that our weights can be easily transferred to any standard Transformer implementation.”; (page 6 section 4.1 “downstream”) “we evaluate the pre-trained VATT models on 4 major downstream tasks[*Examiner notes: inputting task type] using a total of 10 datasets.”
wherein the target model comprises a feature extraction module,
(page 4 section 3) “We feed each modality to a tokenization layer, where the raw input is projected to an embedding vector followed by a Transformer[*Examiner notes: feature extraction module]. There are two major settings: 1) The backbone Transformers are separate and have specific weights for each modality, and 2) The Transformers share weights, namely, there is a single backbone Transformer applied to any of the modalities.”
a plurality of object segmentation modules, different object segmentation modules correspond to different object types,
(page 4 section 3.1) “The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer[*Examiner notes: plurality of object segmentation modules] that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers.”
and a plurality of task processing modules
and different task processing modules correspond to different task types.
(page 17 section A.1.2)
PNG
media_image3.png
416
608
media_image3.png
Greyscale
Regarding Claim 11
Akbari teaches:
An object processing apparatus, comprising: an object obtaining module, configured to obtain a target object to be processed;
(page 1 abstract) “Specifically, our Video Audio-Text Transformer (VATT) takes raw signals as inputs[*Examiner notes: target object to be processed] and extracts multi modal representations that are rich enough to benefit a variety of downstream tasks.”
a first determining module, configured to determine an object type of the target object;
(page 4 section 3.1) “VATT operates on raw signals. The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers.”; [*Examiner notes: The object type is determined when it is input to the modality-specific tokenization layer]
a second determining module, configured to determine a task type corresponding to a target task for processing the target object;
(page 6 section 4.1 “Downstream”) “We evaluate the pre-trained Transformers on a variety of downstream tasks: image classification, video action recognition, audio event classification, and zero-shot text-to-video retrieval.”
an object processing module, configured to input the target object, the object type and the task type into a pre-generated target model to obtain a target result output by the target model;
(page 4 section 3.2) “For simplicity, we adopt the most established Transformer architecture[*Examiner notes: pre-generated model] [23], which has been widely used in NLP. Similar to ViT [25], we do not tweak the architecture so that our weights can be easily transferred to any standard Transformer implementation.”; (page 6 section 4.1 “downstream”) “we evaluate the pre-trained VATT models on 4 major downstream tasks[*Examiner notes: inputting task type] using a total of 10 datasets.”
wherein the target model comprises a feature extraction module,
(page 4 section 3) “We feed each modality to a tokenization layer, where the raw input is projected to an embedding vector followed by a Transformer[*Examiner notes: feature extraction module]. There are two major settings: 1) The backbone Transformers are separate and have specific weights for each modality, and 2) The Transformers share weights, namely, there is a single backbone Transformer applied to any of the modalities.”
a plurality of object segmentation modules […], different object segmentation modules correspond to different object types,
(page 4 section 3.1) “The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer[*Examiner notes: plurality of object segmentation modules] that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers.”
and a plurality of task processing modules
and different task processing modules correspond to different task types.
(page 17 section A.1.2)
PNG
media_image3.png
416
608
media_image3.png
Greyscale
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-5 and 14-18 are rejected under 35 U.S.C. 103 as being unpatentable over Akbari in view of Wu et al. “UNDERSTANDING AND IMPROVING INFORMATION TRANSFER IN MULTI-TASK LEARNING” herein referred to as Wu.
Regarding Claim 2
Akbari teaches:
The method according to claim 1
(see rejection of claim 1)
wherein the inputting the target object, the object type and the task type into the pre-generated target model to obtain the target result output by the target model comprises: inputting the target object into a target segmentation module to obtain a plurality of first object features after segmentation, wherein the target segmentation module comprises an object segmentation module corresponding to the object type;
(page 4 section 3.1) “VATT operates on raw signals. The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers. Besides, each modality has its own positional encoding, which injects the order of tokens into Transformers [88].”
inputting the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module;
(page 5 below equation 2) “where xn is the input patches sequence and xAGG is the learnable embedding of a special aggregation token whose corresponding output in the Transformer[*Examiner notes: ] (z0 out) is used as the aggregated representation for the entire input sequence. This will be later used for classification and common space mapping”
Akbari does not explicitly teach:
inputting the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type.
However, Wu teaches:
inputting the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type.
(page 1 abstract) “We investigate multi-task learning approaches that use a shared feature representation for all tasks. To better understand the transfer of task information, we study an architecture with a shared module for all tasks and a separate output module for each task[*Examiner notes: task processing module corresponding to task type].”
PNG
media_image4.png
243
205
media_image4.png
Greyscale
Akbari, Wu, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari by using task processing modules corresponding to task types as taught by Wu because (Wu page 1 last paragraph) “Our motivating observation is that in addition to model similarity which affects the type of interference, task data similarity plays a second-order effect after controlling model similarity” and (Wu page 1 abstract) “Inspired by the theoretical insights, we show that aligning tasks’ embedding layers leads to performance gains for multi-task training and transfer learning on the GLUE benchmark and sentiment analysis tasks; for example, we obtain a 2.35% GLUE score average improvement on 5 GLUE tasks over BERTLARGE using our alignment method.”
Regarding Claim 3
Akbari in view of Wu teaches:
The method according to claim 2
(see rejection of claim 2)
Akbari further teaches:
wherein the feature extraction module comprises a shared layer and candidate adaptation layers, each of the candidate adaptation layers corresponds to a different task type;
[*Examiner notes: The candidate adaptations are the linear projections while the segmentation modules are the tokenizations]; (page 3 paragraph 3) “To move one step forward, we challenge the Transformers in VATT by a seemingly too strong constraint: sharing weights among the video, audio, and text modalities. The idea is to test whether there exists a single, general-purpose model for all the modalities — of course, they still have their own layers of tokenization and linear projection.”; Figure 1
PNG
media_image5.png
370
947
media_image5.png
Greyscale
the inputting the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module comprises: performing feature extraction on the plurality of first object features according to the shared layer and a target adaptation layer to obtain the second object feature,
(page 4 section 3) “We feed each modality to a tokenization layer, where the raw input is projected[*Examiner notes: target adaptation layer] to an embedding vector followed by a Transformer[*Examiner notes: shared layer].”
wherein the target adaptation layer is a candidate adaptation layer corresponding to the task type.
(page 3 paragraph 3) “To move one step forward, we challenge the Transformers in VATT by a seemingly too strong constraint: sharing weights among the video, audio, and text modalities. The idea is to test whether there exists a single, general-purpose model for all the modalities — of course, they still have their own layers of tokenization and linear projection.”; Figure 1
Regarding Claim 4
Akbari teaches:
The method according to claim 1
(see rejection of claim 1)
Akbari does not explicitly teach:
wherein the target model is generated by: obtaining a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types;
determining a second sample set according to the plurality of first sample sets; training a multimodal model according to the second sample set to obtain the target model.
However, Wu teaches:
wherein the target model is generated by: obtaining a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types;
(page 3 paragraph 1) “We are given k tasks[*Examiner notes: different first sample sets correspond to different task types]. Let mi denote the number of data samples of task i. For task i, let Xi ∈ Rmi× d denote its covariates[*Examiner notes: plurality of sample objects] and let yi ∈ Rmi denote its labels[*Examiner notes: sample result corresponding to each object], where d is the dimension of the data”
determining a second sample set according to the plurality of first sample sets; training a multimodal model according to the second sample set to obtain the target model.
(page 3 paragraph 1) “We define the objective of finding an MTL model as minimizing the following equation over B and the Ai’s: [equation 1]”; (page 3 above equation 2) “. In equation 1, all data samples contribute equally. Because of the differences between tasks such as data size, it is natural to re-weight tasks during training [equation 2]”
Akbari, Wu, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari by using the samples and training as taught by Wu because (Wu page 1 abstract) “Inspired by the theoretical insights, we show that aligning tasks’ embedding layers leads to performance gains for multi-task training and transfer learning on the GLUE benchmark and sentiment analysis tasks; for example, we obtain a 2.35% GLUE score average improvement on 5 GLUE tasks over BERTLARGE using our alignment method. We also design an SVD-based task reweighting scheme and show that it improves the robustness of multi-task training on a multi-label image dataset”
Regarding Claim 5
Akbari in view of Wu teaches:
The method according to claim 4
(see rejection of claim 4)
And Wu further teaches:
wherein the determining a second sample set according to the plurality of first sample sets comprises: taking the plurality of first sample sets as the second sample set;
(page 3 below equation 1) “. In equation 1, all data samples contribute equally.”; Equation 1
PNG
media_image6.png
53
447
media_image6.png
Greyscale
or sampling the plurality of first sample set according to task types to obtain the second sample set.
(page 3 above equation 2) “Because of the differences between tasks such as data size, it is natural to re-weight tasks during training:”; Equation 2
PNG
media_image7.png
38
431
media_image7.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari with Wu for the same reasons given in claim 4 above.
Regarding Claim 14
Claim 14 is nearly identical to claim 4. The only difference is that claim 14 is dependent upon claim 2 instead of claim 1. Therefore, the same rejection and rationale applies with reference to claim 2 instead of claim 1.
Regarding Claim 15
Claim 15 is nearly identical to claim 4. The only difference is that claim 15 is dependent upon claim 3 instead of claim 1. Therefore, the same rejection and rationale applies with reference to claim 3 instead of claim 1.
Regarding Claim 16
Akbari teaches:
The object processing apparatus according to claim 11
(see rejection of claim 11)
wherein the object processing module is further configured to: input the target object into a target segmentation module to obtain a plurality of first object features after segmentation, wherein the target segmentation module comprises an object segmentation module corresponding to the object type;
(page 4 section 3.1) “VATT operates on raw signals. The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers. Besides, each modality has its own positional encoding, which injects the order of tokens into Transformers [88].”
input the plurality of first object features into the feature extraction module to obtain a second object feature output by the feature extraction module;
(page 5 below equation 2) “where xn is the input patches sequence and xAGG is the learnable embedding of a special aggregation token whose corresponding output in the Transformer[*Examiner notes: ] (z0 out) is used as the aggregated representation for the entire input sequence. This will be later used for classification and common space mapping”
Akbari does not explicitly teach:
input the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type.
However, Wu teaches:
input the second object feature into a target processing module to obtain the target result output by the target processing module, wherein the target processing module comprises a task processing module corresponding to the task type.
(page 1 abstract) “We investigate multi-task learning approaches that use a shared feature representation for all tasks. To better understand the transfer of task information, we study an architecture with a shared module for all tasks and a separate output module for each task[*Examiner notes: task processing module corresponding to task type].”
PNG
media_image4.png
243
205
media_image4.png
Greyscale
Akbari, Wu, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari by using task processing modules corresponding to task types as taught by Wu because (Wu page 1 last paragraph) “Our motivating observation is that in addition to model similarity which affects the type of interference, task data similarity plays a second-order effect after controlling model similarity” and (Wu page 1 abstract) “Inspired by the theoretical insights, we show that aligning tasks’ embedding layers leads to performance gains for multi-task training and transfer learning on the GLUE benchmark and sentiment analysis tasks; for example, we obtain a 2.35% GLUE score average improvement on 5 GLUE tasks over BERTLARGE using our alignment method.”
Regarding Claim 17
Akbari teaches:
The object processing apparatus according to claim 11,
(see rejection of claim 11)
Akbari does not explicitly teach:
further comprising a model generation module, configured to: obtain a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types;
determine a second sample set according to the plurality of first sample sets; train a multimodal model according to the second sample set to obtain the target model.
However, Wu teaches:
further comprising a model generation module, configured to: obtain a plurality of first sample sets, wherein each of the plurality of first sample sets comprises a plurality of sample objects and a sample result corresponding to each of the plurality of sample objects, and different first sample sets correspond to different task types;
(page 3 paragraph 1) “We are given k tasks[*Examiner notes: different first sample sets correspond to different task types]. Let mi denote the number of data samples of task i. For task i, let Xi ∈ Rmi× d denote its covariates[*Examiner notes: plurality of sample objects] and let yi ∈ Rmi denote its labels[*Examiner notes: sample result corresponding to each object], where d is the dimension of the data”
determine a second sample set according to the plurality of first sample sets; train a multimodal model according to the second sample set to obtain the target model.
(page 3 paragraph 1) “We define the objective of finding an MTL model as minimizing the following equation over B and the Ai’s: [equation 1]”; (page 3 above equation 2) “. In equation 1, all data samples contribute equally. Because of the differences between tasks such as data size, it is natural to re-weight tasks during training [equation 2]”
Akbari, Wu, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari by using the samples and training as taught by Wu because (Wu page 1 abstract) “Inspired by the theoretical insights, we show that aligning tasks’ embedding layers leads to performance gains for multi-task training and transfer learning on the GLUE benchmark and sentiment analysis tasks; for example, we obtain a 2.35% GLUE score average improvement on 5 GLUE tasks over BERTLARGE using our alignment method. We also design an SVD-based task reweighting scheme and show that it improves the robustness of multi-task training on a multi-label image dataset”
Regarding Claim 18Akbari in view of Wu teaches:
The object processing apparatus according to claim 17
(see rejection of claim 17)
And Wu further teaches:
wherein the model generation module is further configured to: take the plurality of first sample sets as the second sample set;
(page 3 below equation 1) “. In equation 1, all data samples contribute equally.”; Equation 1
PNG
media_image6.png
53
447
media_image6.png
Greyscale
or sample the plurality of first sample set according to task types to obtain the second sample set.
(page 3 above equation 2) “Because of the differences between tasks such as data size, it is natural to re-weight tasks during training:”; Equation 2
PNG
media_image7.png
38
431
media_image7.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari with Wu for the same reasons given in claim 17 above.
Claims 6 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Akbari in view of Wu, and further in view of NPL reference Liu et al. “A Weighted-resampling based Transfer Learning Algorithm” herein referred to as Liu.
Regarding Claim 6
Akbari in view of Wu teaches:
The method according to claim 5
(see rejection of claim 5)
And Wu further teaches:
wherein the sampling the plurality of first sample set according to task types to obtain the second sample set comprises: determining sampling weights according to the task types;
(Page 3 above equation 2) “Because of the differences between tasks such as data size, it is natural to re-weight tasks during training:”; Equation 2
PNG
media_image7.png
38
431
media_image7.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari with Wu for the same reasons given in claim 4 above.
Akbari in view of Wu does not explicitly teach:
sampling each of the plurality of first sample sets according to the sampling weights to obtain a third sample set corresponding to the each of the plurality of first sample sets; determining the second sample set according to the third sample set.
However, Liu teaches:
sampling each of the plurality of first sample sets according to the sampling weights to obtain a third sample set corresponding to the each of the plurality of first sample sets; determining the second sample set according to the third sample set.
(page 186 column 2 last paragraph) “The main idea of TrResampling algorithm as follow. In each iteration round, a new source training data set is created by weighted-resampling from the original source data sets, then the labeled data in the target data set are added to the new source training data set as the training data set. As a result, the data assigned higher weights in the source data set will receive more emphasis.”
Akbari, Wu, Liu, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari in view of Wu by using the weighted data sampling for training data as taught by Liu because (Liu page 188 column 2 section V “conclusion”) “Experimental results show that our algorithm has outstanding performance comparing with TrAdaBoost on many UCI datasets, and obviously improve the classification accuracy based on the base learners”
Regarding Claim 19
The object processing apparatus according to claim 18
(see rejection of claim 18)
And Wu further teaches:
wherein the model generation module is further configured to: determine sampling weights according to the task types;
(Page 3 above equation 2) “Because of the differences between tasks such as data size, it is natural to re-weight tasks during training:”; Equation 2
PNG
media_image7.png
38
431
media_image7.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari with Wu for the same reasons given in claim 17 above.
Akbari in view of Wu does not explicitly teach:
sample each of the plurality of first sample sets according to the sampling weights to obtain a third sample set corresponding to the each of the plurality of first sample sets; determine the second sample set according to the third sample set.
However, Liu teaches:
sample each of the plurality of first sample sets according to the sampling weights to obtain a third sample set corresponding to the each of the plurality of first sample sets; determine the second sample set according to the third sample set.
(page 186 column 2 last paragraph) “The main idea of TrResampling algorithm as follow. In each iteration round, a new source training data set is created by weighted-resampling from the original source data sets, then the labeled data in the target data set are added to the new source training data set as the training data set. As a result, the data assigned higher weights in the source data set will receive more emphasis.”
Akbari, Wu, Liu, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari in view of Wu by using the weighted data sampling for training data as taught by Liu because (Liu page 188 column 2 section V “conclusion”) “Experimental results show that our algorithm has outstanding performance comparing with TrAdaBoost on many UCI datasets, and obviously improve the classification accuracy based on the base learners”
Claims 7 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Akbari in view of Wu and Liu, and further in view of NPL reference Bingel et al. “Identifying beneficial task relations for multi-task learning in deep neural networks” herein referred to as Bingel.
Regarding Claim 7
Akbari in view of Wu and Liu teaches:
The method according to claim 6
(see rejection of claim 6)
Liu further teaches:
wherein the determining the second sample set according to the third sample set comprises: taking the third sample set as the second sample set;
(page 186 column 2 last paragraph) “The main idea of TrResampling algorithm as follow. In each iteration round, a new source training data set is created by weighted-resampling from the original source data sets, then the labeled data in the target data set are added to the new source training data set as the training data set. As a result, the data assigned higher weights in the source data set will receive more emphasis.”
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari and Wu with Liu for the same reasons given in claim 6 above.
Akbari in view of Wu and Liu does not explicitly teach:
or taking a third sample set with a same task type as a task processing module of the target model as the second sample set.
However, Bingel teaches:
or taking a third sample set with a same task type as a task processing module of the target model as the second sample set.
(page 165 column 2 paragraph 3) “In our experiments below, we consider the following ten NLP tasks, with one dataset for each task”; (page 165 column 2 paragraph 2) “In our MTL setup, a training step consists of uniformly drawing a training task, then sampling a random batch of 32 examples from the task’s training data. Every training step thus works on exactly one task, and optimizes the task-specific projection and the shared parameters using Adadelta.”
Akbari, Wu, Liu, Bingel, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari in view of Wu and Liu by using a sample set with the same task type as the processing module as the second sample set because (Bingel page 167 paragraph 1) “In almost four in five cases, we can predict the outcome of the MTL experiment from the data and the single task experiments, which gives validity to our feature analysis. We also see that the features derived from the single task inductions are the most important” and (Bingel page 164 column 2 paragraph 1) “Running MTL experiments on 90 task configurations and comparing their performance to single-task setups, we identify data characteristics and patterns in single-task learning that predict task synergies in deep neural networks.” That is, it’s compelling to run experiments to find out where specializing training to individualized tasks performs better than training using multiple task sources.
Regarding Claim 20
Akbari in view of Wu and Liu teaches:
The method according to claim 19
(see rejection of claim 19)
Liu further teaches:
wherein the model generation module is further configured to: take the third sample set as the second sample set;
(page 186 column 2 last paragraph) “The main idea of TrResampling algorithm as follow. In each iteration round, a new source training data set is created by weighted-resampling from the original source data sets, then the labeled data in the target data set are added to the new source training data set as the training data set. As a result, the data assigned higher weights in the source data set will receive more emphasis.”
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari and Wu with Liu for the same reasons given in claim 19 above.
Akbari in view of Wu and Liu does not explicitly teach:
or take a third sample set with a same task type as a task processing module of the target model as the second sample set.
However, Bingel teaches:
or take a third sample set with a same task type as a task processing module of the target model as the second sample set.
(page 165 column 2 paragraph 3) “In our experiments below, we consider the following ten NLP tasks, with one dataset for each task”; (page 165 column 2 paragraph 2) “In our MTL setup, a training step consists of uniformly drawing a training task, then sampling a random batch of 32 examples from the task’s training data. Every training step thus works on exactly one task, and optimizes the task-specific projection and the shared parameters using Adadelta.”
Akbari, Wu, Liu, Bingel, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari in view of Wu and Liu by using a sample set with the same task type as the processing module as the second sample set because (Bingel page 167 paragraph 1) “In almost four in five cases, we can predict the outcome of the MTL experiment from the data and the single task experiments, which gives validity to our feature analysis. We also see that the features derived from the single task inductions are the most important” and (Bingel page 164 column 2 paragraph 1) “Running MTL experiments on 90 task configurations and comparing their performance to single-task setups, we identify data characteristics and patterns in single-task learning that predict task synergies in deep neural networks.” That is, it’s compelling to run experiments to find out where specializing training to individualized tasks performs better than training using multiple task sources.
Claims 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over Akbari in view of Wu, and further in view of NPL reference Tang et al. “Cyclic Autoencoder for Multimodal Data Alignment Using Custom Datasets” herein referred to as Tang.
Regarding Claim 8
Akbari in view of Wu teaches:
The method according to claim 4
(see rejection of claim 4)
And Wu further teaches:
wherein the model training step comprises: obtaining a task loss value of each task processing module of the multimodal model according to the second sample set, wherein the task loss value is used to characterize a difference between a prediction result output by the task processing module and the sample result;
(page 3 paragraph 1) “We define the objective of finding an MTL model as minimizing the following equation over B and the Ai’s: […] Because of the differences between tasks such as data size, it is natural to re-weight tasks during training:”
PNG
media_image8.png
78
491
media_image8.png
Greyscale
calculating a comprehensive loss value according to the task loss values and the task weights of the plurality of task processing modules;
(Equation 2)
PNG
media_image9.png
105
651
media_image9.png
Greyscale
updating, in a case where the multi-modal model is not determined to meet the preset iteration stopping condition according to the comprehensive loss value, parameters of the multi-modal model to obtain a trained multimodal model, and taking the trained multimodal model as a new multimodal model
(page 5 algorithm 1 step 2) “Minimize ˆf by alternatively applying a gradient descent update on Ai and Ri, given a sampled data batch from task i.”
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari with Wu for the same reasons given in claim 4 above.
Akbari in view of Wu does not explicitly teach:
wherein the training a multimodal model according to the second sample set to obtain the target model comprises: performing a model training step cyclically according to the second sample set, until the trained multimodal model is determined to meet a preset iteration stopping condition, and taking the trained multimodal model as the target model;
However, Tang teaches:
wherein the training a multimodal model according to the second sample set to obtain the target model comprises: performing a model training step cyclically according to the second sample set, until the trained multimodal model is determined to meet a preset iteration stopping condition, and taking the trained multimodal model as the target model;
(page 43 paragraph 2) “To be able to stop the training at the appropriate time, we set a termination condition for the model's training. Training is stopped when the accuracy of the model in the validation set reaches 98% or more.”; (page 49 paragraph 2) “We trained the cyclic autoencoder for a total of 450 iterations over 20 rounds using the above experimental setup, and the training and validation results recorded during the training process are shown in Fig. 11 below”
Akbari, Wu, Tang, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari in view of Wu by using the cyclic training and training stop condition as taught by Tang because (Tang page 43 paragraph 2) “To be able to stop the training at the appropriate time, we set a termination condition for the model's training
Regarding Claim 9
Akbari in view of Wu and Tang:
The method according to claim 8
(see rejection of claim 8)
Akbari further teaches
wherein the obtaining a task loss value of each task processing module of the multimodal model according to the second sample set comprises: inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation, wherein the sample segmentation module comprises an object segmentation module corresponding to an object type of the sample object;
(page 4 section 3.1) “VATT operates on raw signals. The vision-modality input consists of 3-channel RGB pixels of video frames, the audio input is in the form of air density amplitudes (waveforms), and the text input is a sequence of words. We first define a modality-specific tokenization layer that takes as input the raw signals and returns a sequence of vectors to be fed to the Transformers. Besides, each modality has its own positional encoding, which injects the order of tokens into Transformers [88].”
inputting the plurality of the first sample object features into the feature extraction module to obtain a second sample object feature output by the feature extraction module;
(page 4 section 3) “We feed each modality to a tokenization layer, where the raw input is projected to an embedding vector followed by a Transformer.”
Wu further teaches:
inputting the second sample object feature into a sample processing module to obtain a prediction result output by the sample processing module, wherein the sample processing module comprises a task processing module corresponding to the task type;
(page 1 abstract) “We investigate multi-task learning approaches that use a shared feature representation for all tasks. To better understand the transfer of task information, we study an architecture with a shared module for all tasks and a separate output module for each task[*Examiner notes: task processing module corresponding to task type].”
PNG
media_image4.png
243
205
media_image4.png
Greyscale
obtaining the task loss value of each task processing module according to the prediction result and the sample result.
(page 3 above equation 1) “We define the objective of finding an MTL model as minimizing the following equation over B and the Ai’s: […] where L is a loss function such as the squared loss. […] Because of the differences between tasks such as data size, it is natural to re-weight tasks during training”; equation 2
PNG
media_image10.png
34
290
media_image10.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to combine Akbari and Tang with Wu for the same reasons given in claim 4 above.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Akbari in view of Wu and Tang, and further in view of NPL reference Deng et al. “FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting” herein referred to as Deng.
Regarding Claim 10
Akbari in view of Wu and Tang teaches:
The method according to claim 9
(see rejection of claim 9)
Akbari in view of Wu and Tang does not explicitly teach:
wherein the inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation comprises: segmenting, in a case where the object type of the sample object is image or audio, to obtain a plurality of sub-sample objects, wherein an overlapping region exists between sub-sample objects that are adjacent to each other; determining the plurality of first sample object features according to the plurality of sub-sample objects.
However, Deng teaches:
wherein the inputting the sample object into a sample segmentation module to obtain a plurality of first sample object features after segmentation comprises: segmenting, in a case where the object type of the sample object is image or audio, to obtain a plurality of sub-sample objects, wherein an overlapping region exists between sub-sample objects that are adjacent to each other; determining the plurality of first sample object features according to the plurality of sub-sample objects.
(page 14040 abstract) “On the contrary, the soft composition operates by stitching different patches into a whole feature map where pixels in overlapping regions are summed up. These two modules are first used in tokenization before Transformer layers and de-tokenization after Transformer layers, for effective mapping between tokens and features. Therefore, sub-patch level information interaction is enabled for more effective feature propagation between neighboring patches, resulting in synthesizing vivid content for hole regions in videos.”
Akbari, Wu, Deng, and the instant application are analogous because they are all directed to neural networks and/or transformers.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Akbari in view of Wu by using the overlapping sub-sample objects during segmentation as taught by Deng because (Deng page 14040 abstract) “Therefore, sub-patch level information interaction is enabled for more effective feature propagation between neighboring patches, resulting in synthesizing vivid content for hole regions in videos.”
Claims 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Akbari in view of Grymel et al. (US20230020929A1) herein referred to as Grymel.
Regarding Claim 12
Akbari teaches:
realizes steps of the method according to claim 1.
(see rejection of claim 1)
Akbari does not explicitly teach:
A computer-readable medium, storing a computer program thereon, wherein the computer program, when executed by a processing apparatus
However, Grymel teaches:
A computer-readable medium, storing a computer program thereon, wherein the computer program, when executed by a processing apparatus
(paragraph [0168]) “In some embodiments, the memory 1404 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for deep learning, e.g., the method 1100 described above in conjunction with FIG. 11 or some operations performed by the compute tile described above in conjunction with FIG. 3 (e.g., operations performed by the WCB 320).”
Akbari, Grymel, and the instant application are analogous because they are all directed to machine learning.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Wu by implementing it with the computer-readable medium of Grymel because because (Grymel paragraph [0168]) “In some embodiments, the memory 1404 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for deep learning, e.g., the method 1100 described above”. That is, the computer-readable medium can be used to store and execute computer code related to deep learning.
Regarding Claim 13
Akbari teaches:
realize steps of the method according to claim 1
(see rejection of claim 1)
Akbari does not explicitly teach:
An electronic device, comprising: a storage apparatus, storing a computer program thereon; a processing apparatus, configured to execute the computer program on the storage apparatus to
However, Grymel teaches:
An electronic device, comprising: a storage apparatus, storing a computer program thereon; a processing apparatus, configured to execute the computer program on the storage apparatus to
(paragraph [0168]) “The computing device 1400 may include a processing device 1402 (e.g., one or more processing devices). The processing device 1402 processes electronic data from registers and/or memory to transform that electronic data into other electronic data that may be stored in registers and/or memory.”
Akbari, Grymel, and the instant application are analogous because they are all directed to machine learning.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention to modify the task processing transformer taught by Wu by implementing it with the computer of Grymel because (Grymel paragraph [0168]) “In some embodiments, the memory 1404 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for deep learning, e.g., the method 1100 described above”. That is, the computer can be used to store and execute computer code related to deep learning.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Hu et al. “UniT: Multimodal Multitask Learning with a Unified Transformer” teaches multimodal and multitask learning using transformers.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ezra J Baker whose telephone number is (703)756-1087. The examiner can normally be reached Monday - Friday 10:00 am - 8:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/E.J.B./Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126