Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 5/16/2024 and 6/17/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Status of Claims
The present application is being examined under the claims filed on 5/16/2024.
Claims 1-20 are rejected.
Claims 1-20 are pending.
Specification
The specification filed on 5/16/2024 is acceptable for examination purposes.
Drawings
The drawings filed on 5/16/2024 are acceptable for examination purposes.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) are: “a control unit configured to control operations of the first inference model and the second inference model” and “a determination unit configured to determine whether or not each of the plurality of data items is usable for the inference processing” in claim 1.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-15 and 17-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as failing to set forth the subject matter which the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the applicant regards as the invention.
Claim limitations “a control unit configured to control operations of the first inference model and the second inference model” and “a determination unit configured to determine whether or not each of the plurality of data items is usable for the inference processing” [claim 1] invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. In particular, the specification, at best, describes “a control unit configured to control operations of the first inference model and the second inference model” and “a determination unit configured to determine whether or not each of the plurality of data items is usable for the inference processing” (Spec. paragraph [0006]) and does not provide any specific hardware or structure to support the claimed control unit and determination unit. Thus, the disclosure provides no association between the structure and the function can be found in the specification. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph.
In reference to dependent claims 2-15 and 17-19, claims 2-15 and 17-19 do not cure the deficiencies noted in the rejection of independent claim 1. Therefore, these claims are rejected under the same rationale as claim 1.
For the purpose of this examination, limitations “control unit” is interpreted as software module loaded in a controller 103 and “determination unit” is interpreted as software module loaded in a determination section 102. See MPEP 2173.06(I).
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Claim 1,
Step 1: Claim 1 is a control device claim. Therefore, Claims 1-15 are directed to a machine.
Step 2A Prong 1:
[a determination unit configured to] determine whether or not each of the plurality of data items is usable for the inference processing (mental process – determining whether or not each of the plurality of data items is usable for the inference processing may be performed manually by a user with the aid of pen and paper by observing/analyzing the plurality of data items and using a judgement to determine whether each of the plurality of data items is usable for the inference processing or not. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a control unit configured to control operations of the first inference model and the second inference model (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a determination unit configured to (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
wherein the control unit controls, in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values, and the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers, and controls, in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by the determination unit to be usable for the inference processing as an input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as a first inference model, a second inference model, a control unit and a determination unit.
Additional Elements:
a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a control unit configured to control operations of the first inference model and the second inference model (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a determination unit configured to (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
wherein the control unit controls, in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values, and the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers, and controls, in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by the determination unit to be usable for the inference processing as an input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-15. The additional limitations of the dependent claims are addressed below.
Regarding Claim 2,
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 2 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the control unit performs control to input a predetermined fixed value to the connection layer in place of an output from the second inference model (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the control unit performs control to input a predetermined fixed value to the connection layer in place of an output from the second inference model (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Regarding Claim 3,
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 3 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the control unit performs control to input a predetermined fixed value to the connection layer in place of an output from a layer to which the data determined by the determination unit to be unusable for the inference processing is input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the control unit performs control to input a predetermined fixed value to the connection layer in place of an output from a layer to which the data determined by the determination unit to be unusable for the inference processing is input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Regarding Claim 4,
Step 2A Prong 1:
[wherein the determination unit calculates an evaluation value of each of the plurality of data items], and determines whether or not each data item is usable for the inference processing based on a result of comparison between the calculated evaluation value and a predetermined threshold value (mental process – determining whether or not each data item is usable for the inference processing based on a result of comparison between the calculated evaluation value and a predetermined threshold value may be performed manually by a user with the aid of pen and paper by observing/analyzing a result of comparison between the calculated evaluation value and a predetermined threshold value, and using a judgement to determine whether or not each data is usable for the inference processing. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the determination unit calculates an evaluation value of each of the plurality of data items (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the determination unit calculates an evaluation value of each of the plurality of data items (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Regarding Claim 5,
Step 2A Prong 1:
See the rejection of Claim 4 above, which Claim 5 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the plurality of data items are image data and sound data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the plurality of data items are image data and sound data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 6,
Step 2A Prong 1:
See the rejection of Claim 5 above, which Claim 6 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the evaluation value is a value related to one of a luminance of each image data item, a blur amount of the image data item, and a high-sensitivity noise amount of the image data item (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the evaluation value is a value related to one of a luminance of each image data item, a blur amount of the image data item, and a high-sensitivity noise amount of the image data item (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 7,
Step 2A Prong 1:
See the rejection of Claim 5 above, which Claim 7 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the evaluation value is a value related to a noise amount of a feature component of the sound data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the evaluation value is a value related to a noise amount of a feature component of the sound data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 8,
Step 2A Prong 1:
See the rejection of Claim 4 above, which Claim 8 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the plurality of data items are a plurality of image data items obtained by capturing an image of a person at different angles (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the plurality of data items are a plurality of image data items obtained by capturing an image of a person at different angles (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 9,
Step 2A Prong 1:
See the rejection of Claim 8 above, which Claim 9 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the evaluation value is a value related to an orientation of a face of a person whose image appears in the image data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the evaluation value is a value related to an orientation of a face of a person whose image appears in the image data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 10,
Step 2A Prong 1:
See the rejection of Claim 4 above, which Claim 10 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the plurality of data items include image data of interest and reference image data for performing noise reduction for eliminating noise from the image data of interest (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the plurality of data items include image data of interest and reference image data for performing noise reduction for eliminating noise from the image data of interest (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 11,
Step 2A Prong 1:
See the rejection of Claim 10 above, which Claim 11 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the evaluation value is a value related to a difference between the image data of interest and the reference image data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the evaluation value is a value related to a difference between the image data of interest and the reference image data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 12,
Step 2A Prong 1:
[wherein the determination unit] determines whether or not each image data item is usable for the inference processing, based on image-capturing conditions at the time of capturing the image data items (mental process – determining whether or not each image data item is usable for the inference processing, based on image-capturing conditions at the time of capturing the image data items may be performed manually by a user with the aid of pen and paper by observing/analyzing the image-capturing conditions of the image data items and using a judgement to determine whether or not each image data item is usable for the inference processing. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the plurality of data items are image data items obtained by capturing an image of an object (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
wherein the determination unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the plurality of data items are image data items obtained by capturing an image of an object (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
wherein the determination unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Regarding Claim 13,
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 13 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the plurality of data items are two data items, and wherein the first inference model includes two layers to which the two data items are input, respectively (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the plurality of data items are two data items, and wherein the first inference model includes two layers to which the two data items are input, respectively (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 14,
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 14 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the plurality of data items are three or more data items, and wherein the first inference model includes three or more layers to which the plurality of data items are input, respectively (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the plurality of data items are three or more data items, and wherein the first inference model includes three or more layers to which the plurality of data items are input, respectively (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 15,
Step 2A Prong 1:
[wherein the determination unit] calculates evaluation values of the plurality of data items, respectively, and sets weighting values to the plurality of data items based on the calculated evaluation values, respectively (mental process - calculating evaluation values of the plurality of data items, respectively, and setting weighting values to the plurality of data items based on the calculated evaluation values, respectively may be performed manually by a user with the aid of pen and paper by observing/analyzing the plurality of data items, calculating the evaluation values to set the weighting values to the plurality of data items. See MPEP 2106.04(a)(2)(III)(C).)
[wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing], [the control unit] determines operation resources of the second inference model based on the weighting values (mental process - determining operation resources of the second inference model based on the weighting values may be performed manually by a user with the aid of pen and paper by observing/analyzing the second inference model and the weighting values, and using a judgement to determine operation resources of the second inference model. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein the determination unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
the control unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein the determination unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
the control unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Regarding Claim 16,
Step 1: Claim 16 is a method claim. Therefore, Claim 16 is directed to a process.
Step 2A Prong 1:
determining whether or not each of the plurality of data items is usable for the inference processing (mental process – determining whether or not each of the plurality of data items is usable for the inference processing may be performed manually by a user with the aid of pen and paper by observing/analyzing the plurality of data items and using a judgement to determine whether or not each of the plurality of data items is usable for the inference processing. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
controlling an operation of a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values, and an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values, and the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as a first inference model, a second inference model, a control unit and a determination unit.
Additional Elements:
controlling an operation of a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values, and an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values, and the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
For the reasons above, Claim 16 is rejected as being directed to an abstract idea without significantly more.
Regarding Claim 17,
Step 1: Claim 17 is a method claim. Therefore, Claims 17-19 are directed to a process.
Step 2A Prong 1:
performing first learning using a plurality of learning data items which are a plurality of associated learning data items and are determined [by the determination unit] to be usable for the inference processing (mental process – performing first learning using a plurality of learning data items which are a plurality of associated learning data items and are determined to be usable for the inference processing may be performed manually by a user with the aid of pen and paper by observing/analyzing a plurality of learning data items and using a judgement to determine to be usable for the inference processing. See MPEP 2106.04(a)(2)(III)(C).)
performing second learning using learning data out of a plurality of learning data items related to each other, which is determined [by the determination unit] to be usable for the inference processing, [a second inference model which operates using the learning data determined to be usable for the reference processing as an input, and the first inference model which performs an operation by a layer using the learning data determined to be usable for the reference processing as an input, and operates by inputting an output from the layer and an output from the second reference model to the connection layer] (mental process – performing second learning using learning data out of a plurality of learning data items related to each other, which is determined to be usable for the inference processing may be performed manually by a user with the aid of pen and paper by observing/analyzing a plurality of learning data items and using a judgement to determine to be usable for the inference processing. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
by the determination unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a second inference model which operates using the learning data determined to be usable for the reference processing as an input, and the first inference model which performs an operation by a layer using the learning data determined to be usable for the reference processing as an input, and operates by inputting an output from the layer and an output from the second reference model to the connection layer (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as a first inference model, a second inference model, a control unit and a determination unit.
Additional Elements:
by the determination unit (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
a second inference model which operates using the learning data determined to be usable for the reference processing as an input, and the first inference model which performs an operation by a layer using the learning data determined to be usable for the reference processing as an input, and operates by inputting an output from the layer and an output from the second reference model to the connection layer (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
For the reasons above, Claim 17 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 18-19. The additional limitations of the dependent claims are addressed below.
Regarding Claim 18,
Step 2A Prong 1:
See the rejection of Claim 17 above, which Claim 18 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein in the first learning, learning is performed by inputting a predetermined fixed value to the connection layer in place of the output from the second inference model (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein in the first learning, learning is performed by inputting a predetermined fixed value to the connection layer in place of the output from the second inference model (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 19,
Step 2A Prong 1:
See the rejection of Claim 17 above, which Claim 19 depends on.
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
wherein in the second learning, learning is performed by inputting a predetermined fixed value to the connection layer in place of an output from a layer of the plurality of layers, which uses learning data determined by the determination unit to be unusable for the inference processing (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional Elements:
wherein in the second learning, learning is performed by inputting a predetermined fixed value to the connection layer in place of an output from a layer of the plurality of layers, which uses learning data determined by the determination unit to be unusable for the inference processing (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).)
Regarding Claim 20,
Step 1: Claim 20 is a non-transitory computer-readable storage medium claim. Therefore, Claim 20 is directed to a manufacture.
Step 2A Prong 1:
determining whether or not each of the plurality of data items is usable for the inference processing (mental process – determining whether or not each of the plurality of data items is usable for the inference processing may be performed manually by a user with the aid of pen and paper by observing/analyzing the plurality of data items and using a judgement to determine whether or not each of the plurality of data items is usable for the inference processing. See MPEP 2106.04(a)(2)(III)(C).)
Step 2A Prong 2: The judicial exceptions are not integrated into a practical application.
Additional Elements:
controlling an operation of a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values, and an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values, and the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as a first inference model, a second inference model, a control unit and a determination unit.
Additional Elements:
controlling an operation of a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values, and an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values, and the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5 and 13-20 are rejected under 35 U.S.C. 103 as being unpatentable over Havaei et al. (“HeMIS: Hetero-Modal Image Segmentation”) (hereinafter Havaei), in view of Garg et al. (US 12361679 B1) (hereinafter Garg), and further in view of Yang et al. (US 20220405537 A1) (hereinafter Yang).
Regarding Claim 1,
Havaei teaches:
“A control device that performs inference processing using a plurality of data items related to each other as inputs, comprising:” (preamble)
“a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values” (Havaei, Fig. 1,
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others.”; Examiner’s note: a first inference model (i.e. modality-specific convolutional neural network) that is formed by a plurality of layers to which the plurality of data items are input (i.e. modality-specific convolutional layers), respectively, for extraction of feature values of input data items (i.e. extraction of feature maps), and a connection layer (i.e. convolutional layer in front end) which outputs output data as an inference result (i.e. classifications outputs) based on the extracted feature values (i.e. extracted feature maps) is taught.)
“a determination unit configured to determine whether or not each of the plurality of data items is usable for the inference processing” (Havaei, Fig. 1 and Section 1, “This approach presents the advantage of being robust to any combinatorial subset of available modalities provided as input, without the need to learn a combinatorial number of imputation models.”; Havaei, Section 2, “Here, we make the HeMIS architecture robust to missing modalities by randomly dropping any number for a given training example […] we start randomly dropping modalities, ensuring a higher probability of dropping zero or one modality only.”; Examiner’s note: a determination unit (i.e. abstraction layer) configured to determine whether or not each of the plurality of data items is usable for the inference processing (i.e. randomly dropping modalities for missing modalities and any combinatorial subset of available modalities provided as input teach input is usable for the inference processing) is taught. See Fig. 1 in the above limitation.)
wherein the control unit controls, in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (Havaei, Fig. 1 and Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others. After a few independent stages, feature maps from all available modalities are merged by computing map-wise statistics such as the mean and the variance, quantities whose expectation does not depend on the number of terms (i.e. modalities) that are provided. After merging, the mean and variance feature maps are concatenated and fed into a final set of convolutional stages to obtain network output.”; Examiner’s note: wherein the control unit controls (i.e. convolutional pipeline), in a case where it is determined by the determination unit (i.e. the convolutional neural network) that all of the plurality of data items are usable for the inference processing (i.e. all available modalities used for the inference processing), the first inference model (i.e. modality-specific convolutional neural network) to output the output data from the connection layer (i.e. output obtained at a final set of convolutional stages) based on a plurality of feature values extracted by the plurality of layers (i.e. extracted feature maps from all available modalities in convolutional layers) is taught.)
Havaei does not explicitly teach:
“a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item”
“a control unit configured to control operations of the first inference model and the second inference model”
wherein the control unit controls, in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values
wherein the control unit controls, in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by the determination unit to be usable for the inference processing as an input
Garg teaches:
“a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item” (Garg, Col. 11, Lines 40-46, “An image set 320 also may be used in the student model 354 as part of teacher student training system 350. Image set 320 may constitute a training set of all of the images used in training the models within teacher student training system 350. Image set 320 may also constitute a subset of images in a larger training set. Images in the training set may be obtained from users.”; Garg, Col. 13, Lines 23-25 and Lines 28-32, “[…] the student model 354 can be trained by processing image set 320 using a visual processor 322 […] The result that is generated by the sets of convolutional layers and max pooling layers may be a matrix of numbers, such as floating-point numbers. The matrix may then be converted to a vector for processing by the set of fully-connected layers.”; Examiner’s note: a second inference model (i.e. the student model) to which any of the plurality of data items are input (i.e. images obtained from users) for extraction of a feature value of each of the any input data item (i.e. a vector for processing by the set of fully-connected layers) is taught.)
“a control unit configured to control operations of the first inference model and the second inference model” (Garg, Col. 11, Lines 60-65, “[…] teacher model 352 may be trained by processing image set 300 using two (or more) processing branches. Each processing branch may process different aspects of the image set 300. In some examples, these could include visual information and separate textual information extracted for the images of the image set 300.”; Garg, Col. 12, Lines 28-31, “[…] the text processor 314 may be implemented using a neural network or a subset of layers of a larger neural network to generate encoded representations of the text for a given input.”; Garg, Col. 13, Lines 23-25, “[…] the student model 354 can be trained by processing image set 320 using a visual processor 322.”; Examiner’s note: a control unit (i.e. a visual processor or a text processor) configured to control operations (i.e. extracting visual information and textual information) of the first inference model (i.e. the teacher model) and the second inference model (i.e. the student model) is taught.)
wherein the control unit controls, in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values (Garg, Col. 3, Lines 21-23, Lines 37-40 and Lines 46-49, “[…] the image set generation system may be configured to determine whether all of the image classes in the preferred image set for the item are present […] the image classification model may be implemented as a multi-branch model that uses different processing branches to process input data in different modalities […] a multi-branch model may include an image-based processing branch for input in an image modality, and a text-based processing branch for input in a text-based modality.”; Garg, Col. 10, Lines 44-46, “If the desired stopping point has been reached at block 208, then at block 210 the text processing branch can be removed from the image classification model.”; Examiner’s note: wherein the control unit controls (i.e. an image-based processing branch and a text-based processing branch), in a case where it is determined by the determination unit (i.e. the image set generation system) that all of the plurality of data items are usable for the inference processing (i.e. all of the image classes in the preferred image set for the item are present), the second inference model not to perform extraction of feature values (i.e. if the desired stopping point reached, the text processing branch is removed from the image classification model which means the image classification model not performing the text processing branch) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item, a control unit to control operations of the first inference model and the second inference model and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. One of ordinary skill would have motivation to combine Havaei and Garg so that “the speed of the system may in some implementations be improved while maintaining an accuracy level” (Garg, Col. 4, Lines 24-26).
Yang teaches:
wherein the control unit controls, in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by the determination unit to be usable for the inference processing as an input (Yang, Paragraphs [0049] and [0050], “The multimodal fusion network 602 receives input modalities 604a, 604b, 604c and extracts features 606a, 606b, 606c from each modality that are feature vectors. The output of the feature extractors 606 is fed into an odd-one-out network 612. The odd-one-out network 612 generates an “inconsistent” modality prediction that is fed to a robust fusion layer 608 along with the output of the feature extractors 606. The robust fusion layer 608 outputs a fused feature vector that is subsequently fed to downstream layers 610 to produce an output […] The network 700 receives features 702 such as output from feature extractor 602a, 602b, and 602c and generates a modality prediction weights 704 such that for each feature channel is an associated modality prediction weight 704a, 704b, and 704c. These modality prediction weights 704a, 704b, and 704c produce an odd-one-out vector that is forwarded to the robust feature fusion layer.”; Yang, Fig. 6,
PNG
media_image2.png
480
656
media_image2.png
Greyscale
; Examiner’s note: wherein the control unit controls (i.e. the multimodal fusion network), in a case where it is determined by the determination unit (i.e. the odd-one-out network) that any of the plurality of data items is unusable for the inference processing (i.e. an odd-one-out vector teaches a data item is unusable for the inference processing), the first inference model to output the output data (i.e. the output of the feature extractors 606) from the connection layer (i.e. feature extractor layer), based on a feature value extracted by each layer (i.e. extracted feature 606a, 606b, 606c) to which data determined by the determination unit (i.e. the feature extractor) to be usable for the inference processing is input (i.e. feature vectors), and a feature value extracted by the second inference model (i.e. a fused feature vector by the robust fusion layer) using the data determined by the determination unit (i.e. the odd-one-out network) to be usable for the inference processing as an input (i.e. a fused feature vector) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item, a control unit to control operations of the first inference model and the second inference model and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. Yang teaches the first inference model and the second inference model for a case where any of data items is unusable for the inference processing. One of ordinary skill would have motivation to combine Havaei, Garg and Yang to “improve[] in robustness of the multimodal machine learning system via training and using an odd-one-out network with a robust fusion layer” (Yang, Paragraph [0002]).
Regarding Claim 2,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 1,” (preamble)
“wherein in a case where it is determined by the determination unit that all of the plurality of data items are usable for the inference processing, the control unit performs control to input a predetermined fixed value to the connection layer in place of an output from the second inference model” (Havaei, Fig. 1 and Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others. After a few independent stages, feature maps from all available modalities are merged by computing map-wise statistics such as the mean and the variance, quantities whose expectation does not depend on the number of terms (i.e. modalities) that are provided.”; Garg, Col. 9, Lines 17-20 and Lines 26-34, “In refining image classification model 160, the refinement may be based on performing additional training (e.g., one or more training epochs) and evaluation using a loss function (e.g., mean squared error loss) […] This process may be repeated in an iterative manner until a desired stopping point is reached. For example, the desired stopping point may correspond to satisfaction of an accuracy metric, exhaustion of a quantity of training time or training iterations, etc.”; Garg, Col. 10, Lines 52-55, “[…] the text processing branch is removed by changing the structure of the machine learning model itself to produce an image classification model with a single processing branch, rather than multiple processing branches”; Examiner’s note: wherein in a case where it is determined by the determination unit (i.e. the convolutional neural network) that all of the plurality of data items are usable for the inference processing (i.e. all available modalities used for the inference processing), the control unit performs control (i.e. refining image classification model) to input a predetermined fixed value (i.e. the desired stopping point corresponding to satisfaction of an accuracy metric) to the connection layer in place of an output from the second inference model (i.e. an image classification model created by changing the structure of the machine learning model itself with a single processing branch, rather than multiple processing branches) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 3,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 1,” (preamble)
“wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the control unit performs control to input a predetermined fixed value to the connection layer in place of an output from a layer to which the data determined by the determination unit to be unusable for the inference processing is input” (Yang, Fig. 7,
PNG
media_image3.png
401
489
media_image3.png
Greyscale
; Yang, Paragraphs [0041], [0049] and [0050], “After the machine-learning algorithm 310 achieves a predetermined performance level (e.g., 100% agreement with the outcomes associated with the training dataset 312), the machine-learning algorithm 310 may be executed using data that is not in the training dataset 312. The trained machine-learning algorithm 310 may be applied to new datasets to generate annotated data […] The multimodal fusion network 602 receives input modalities 604a, 604b, 604c and extracts features 606a, 606b, 606c from each modality that are feature vectors. The output of the feature extractors 606 is fed into an odd-one-out network 612. The odd-one-out network 612 generates an “inconsistent” modality prediction that is fed to a robust fusion layer 608 along with the output of the feature extractors 606. The robust fusion layer 608 outputs a fused feature vector that is subsequently fed to downstream layers 610 to produce an output […] The network 700 receives features 702 such as output from feature extractor 602a, 602b, and 602c and generates a modality prediction weights 704 such that for each feature channel is an associated modality prediction weight 704a, 704b, and 704c. These modality prediction weights 704a, 704b, and 704c produce an odd-one-out vector that is forwarded to the robust feature fusion layer.”; Examiner’s note: wherein in a case where it is determined by the determination unit (i.e. the odd-one-out network) that any of the plurality of data items is unusable for the inference processing (i.e. odd-one-out vector), the control unit performs control (i.e. the multimodal fusion network) to input a predetermined fixed value (i.e. a predetermined performance level (e.g., 100% agreement with the outcomes associated with the training dataset)) to the connection layer (i.e. feature extractor layer) in place of an output from a layer (i.e. generated annotated data) to which the data determined by the determination unit to be unusable for the inference processing is input (i.e. data not meeting the predetermined performance level) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 4,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 1,” (preamble)
“wherein the determination unit calculates an evaluation value of each of the plurality of data items, and determines whether or not each data item is usable for the inference processing based on a result of comparison between the calculated evaluation value and a predetermined threshold value” (Yang, Paragraph [0043], “[…] the machine-learning algorithm 310 may process raw source data 315 and output an indication of a representation of an image […] A machine-learning algorithm 310 may generate a confidence level or factor for each output generated. For example, a confidence value that exceeds a predetermined high-confidence threshold may indicate that the machine-learning algorithm 310 is confident that the identified feature corresponds to the particular feature.”; Examiner’s note: wherein the determination unit (i.e. the machine learning model) calculates an evaluation value of each of the plurality of data items (i.e. confidence value generated by processing raw source data), and determines whether or not each data item is usable for the inference processing based on a result of comparison between the calculated evaluation value and a predetermined threshold value (i.e. a confidence value exceeding a predetermined high-confidence threshold teaches each data item is usable for the inference processing) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 5,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 4,” (preamble)
“wherein the plurality of data items are image data and sound data” (Yang, Paragraph [0027], “The term sensor include an optical, light, imaging, or photon sensor (e.g., a charge-coupled device (CCD), a CMOS active-pixel sensor (APS), infrared sensor (IR), CMOS sensor), an acoustic, sound, or vibration sensor (e.g., microphone, geophone, hydrophone) […]”; Examiner’s note: wherein the plurality of data items (i.e. data items from sensor) are image data (i.e. imaging) and sound data (i.e. sound) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 4 above and applicable herein.
Regarding Claim 13,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 1,” (preamble)
“wherein the plurality of data items are two data items, and wherein the first inference model includes two layers to which the two data items are input, respectively” (Havaei, Fig. 1,
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others […] In our implementation, this consists of two convolutional layers with ReLU”; Examiner’s note: wherein the plurality of data items are two data items (i.e. two modalities), and wherein the first inference model (i.e. modality-specific convolutional neural network) includes two layers to which the two data items are input, respectively (i.e. modality-specific convolutional layers consisting of two convolutional layers and each modality is processed by its own convolutional pipeline as input) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 14,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 1,” (preamble)
“wherein the plurality of data items are three or more data items, and wherein the first inference model includes three or more layers to which the plurality of data items are input, respectively” (Havaei, Fig. 1,
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others.”; Examiner’s note: wherein the plurality of data items are three or more data items (i.e. three or more modalities), and wherein the first inference model (i.e. modality-specific convolutional neural network) includes three or more layers to which the plurality of data items are input, respectively (i.e. modality-specific convolutional layers consisting of three or more convolutional layers and each modality is processed by its own convolutional pipeline as input) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 15,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 14,” (preamble)
“wherein the determination unit calculates an evaluation value of the plurality of data items, respectively, and sets weighting values to the plurality of data items based on the calculated evaluation values, respectively, and wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing, the control unit determines operation resources of the second inference model based on the weighting values” (Yang, Paragraphs [0041], [0043] and [0050], “The machine-learning algorithm 310 may be operated in a learning mode using the training dataset 312 as input. The machine-learning algorithm 310 may be executed over a number of iterations using the data from the training dataset 312. With each iteration, the machine-learning algorithm 310 may update internal weighting factors based on the achieved results […] the machine-learning algorithm 310 may process raw source data 315 and output an indication of a representation of an image […] A machine-learning algorithm 310 may generate a confidence level or factor for each output generated. For example, a confidence value that exceeds a predetermined high-confidence threshold may indicate that the machine-learning algorithm 310 is confident that the identified feature corresponds to the particular feature […] The network 700 receives features 702 such as output from feature extractor 602a, 602b, and 602c and generates a modality prediction weights 704 such that for each feature channel is an associated modality prediction weight 704a, 704b, and 704c. These modality prediction weights 704a, 704b, and 704c produce an odd-one-out vector that is forwarded to the robust feature fusion layer.”; Garg, Col. 10, Lines 56-62, “After the text processing branch is removed at block 210, the visual processing branch can be optionally refined at block 212. Refinement of the classification model may begin with reinitialization of portions of the classifier. Reinitialization may be done by changing the parameters of certain layers (e.g., the weights and biases of the neurons), such as one or more layers of the classifier 120.”; Garg, Col. 11, Lines 22-25, “The teacher student training system 350 may train two or more machine learning models. These machine learning models could include a teacher model 352 and a student model 354.” ; Examiner’s note: wherein the determination unit (i.e. the machine learning model) calculates an evaluation value of the plurality of data items (i.e. confidence value generated by processing raw source data), respectively, and sets weighting values to the plurality of data items based on the calculated evaluation values (i.e. updating internal weighting factors based on the achieved results), respectively, and wherein in a case where it is determined by the determination unit that any of the plurality of data items is unusable for the inference processing (i.e. an odd-one-out vector teaches a data item is unusable for the inference processing), the control unit (i.e. the teacher student training system) determines operation resources of the second inference model based on the weighting values (i.e. changing the parameters of certain layers (e.g., the weights and biases of the neurons) of the student model) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 14 above and applicable herein.
Regarding Claim 16,
Havaei teaches:
“A method of controlling a control device that performs inference processing using a plurality of data items related to each other as inputs, comprising:” (preamble)
controlling an operation of a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values (Havaei, Fig. 1,
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others.”; Examiner’s note: controlling an operation of a first inference model (i.e. modality-specific convolutional neural network) that is formed by a plurality of layers to which the plurality of data items are input (i.e. modality-specific convolutional layers), respectively, for extraction of feature values of input data items (i.e. extraction of feature maps), and a connection layer (i.e. convolutional layer in front end) which outputs output data as an inference result (i.e. classifications outputs) based on the extracted feature values (i.e. extracted feature maps) is taught.)
“determining whether or not each of the plurality of data items is usable for the inference processing” (Havaei, Fig. 1 and Section 1, “This approach presents the advantage of being robust to any combinatorial subset of available modalities provided as input, without the need to learn a combinatorial number of imputation models.”; Havaei, Section 2, “Here, we make the HeMIS architecture robust to missing modalities by randomly dropping any number for a given training example […] we start randomly dropping modalities, ensuring a higher probability of dropping zero or one modality only.”; Examiner’s note: determining whether or not each of the plurality of data items is usable for the inference processing (i.e. randomly dropping modalities for missing modalities and any combinatorial subset of available modalities provided as input teach input is usable for the inference processing) is taught. See Fig. 1 in the above limitation of claim 16.)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (Havaei, Fig. 1 and Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others. After a few independent stages, feature maps from all available modalities are merged by computing map-wise statistics such as the mean and the variance, quantities whose expectation does not depend on the number of terms (i.e. modalities) that are provided. After merging, the mean and variance feature maps are concatenated and fed into a final set of convolutional stages to obtain network output.”; Examiner’s note: controlling (i.e. convolutional pipeline), in a case where it is determined by said determining (i.e. the convolutional neural network) that all of the plurality of data items are usable for the inference processing (i.e. all available modalities used for the inference processing), the first inference model (i.e. modality-specific convolutional neural network) to output the output data from the connection layer (i.e. output obtained at a final set of convolutional stages) based on a plurality of feature values extracted by the plurality of layers (i.e. extracted feature maps from all available modalities in convolutional layers) is taught. See Fig. 1 in the above limitation of claim 16.)
Havaei does not explicitly teach:
controlling an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values
“controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input”
Garg teaches:
controlling an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item (Garg, Col. 11, Lines 40-46, “An image set 320 also may be used in the student model 354 as part of teacher student training system 350. Image set 320 may constitute a training set of all of the images used in training the models within teacher student training system 350. Image set 320 may also constitute a subset of images in a larger training set. Images in the training set may be obtained from users.”; Garg, Col. 13, Lines 23-25 and Lines 28-32, “[…] the student model 354 can be trained by processing image set 320 using a visual processor 322 […] The result that is generated by the sets of convolutional layers and max pooling layers may be a matrix of numbers, such as floating-point numbers. The matrix may then be converted to a vector for processing by the set of fully-connected layers.”; Examiner’s note: controlling an operation of a second inference model (i.e. the student model) using any of the plurality of data items as an input (i.e. images obtained from users) for extraction of a feature value of the input data item (i.e. a vector for processing by the set of fully-connected layers)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values (Garg, Col. 3, Lines 21-23, Lines 37-40 and Lines 46-49, “[…] the image set generation system may be configured to determine whether all of the image classes in the preferred image set for the item are present […] the image classification model may be implemented as a multi-branch model that uses different processing branches to process input data in different modalities […] a multi-branch model may include an image-based processing branch for input in an image modality, and a text-based processing branch for input in a text-based modality.”; Garg, Col. 10, Lines 44-46, “If the desired stopping point has been reached at block 208, then at block 210 the text processing branch can be removed from the image classification model.”; Examiner’s note: controlling (i.e. an image-based processing branch and a text-based processing branch), in a case where it is determined by said determining (i.e. the image set generation system) that all of the plurality of data items are usable for the inference processing (i.e. all of the image classes in the preferred image set for the item are present), the second inference model not to perform extraction of feature values (i.e. if the desired stopping point reached, the text processing branch is removed from the image classification model which means the image classification model not performing the text processing branch) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of the input data item and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. One of ordinary skill would have motivation to combine Havaei and Garg so that “the speed of the system may in some implementations be improved while maintaining an accuracy level” (Garg, Col. 4, Lines 24-26).
Yang teaches:
“controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input” (Yang, Paragraphs [0049] and [0050], “The multimodal fusion network 602 receives input modalities 604a, 604b, 604c and extracts features 606a, 606b, 606c from each modality that are feature vectors. The output of the feature extractors 606 is fed into an odd-one-out network 612. The odd-one-out network 612 generates an “inconsistent” modality prediction that is fed to a robust fusion layer 608 along with the output of the feature extractors 606. The robust fusion layer 608 outputs a fused feature vector that is subsequently fed to downstream layers 610 to produce an output […] The network 700 receives features 702 such as output from feature extractor 602a, 602b, and 602c and generates a modality prediction weights 704 such that for each feature channel is an associated modality prediction weight 704a, 704b, and 704c. These modality prediction weights 704a, 704b, and 704c produce an odd-one-out vector that is forwarded to the robust feature fusion layer.”; Yang, Fig. 6,
PNG
media_image2.png
480
656
media_image2.png
Greyscale
; Examiner’s note: controlling (i.e. the multimodal fusion network), in a case where it is determined by said determining (i.e. the odd-one-out network) that any of the plurality of data items is unusable for the inference processing (i.e. an odd-one-out vector teaches a data item is unusable for the inference processing), the first inference model to output the output data (i.e. the output of the feature extractors 606) from the connection layer (i.e. feature extractor layer), based on a feature value extracted by each layer (i.e. extracted feature 606a, 606b, 606c) to which data determined by the determination unit (i.e. the feature extractor) to be usable for the inference processing is input (i.e. feature vectors), and a feature value extracted by the second inference model (i.e. a fused feature vector by the robust fusion layer) using the data determined by said determining (i.e. the odd-one-out network) to be usable for the inference processing as an input (i.e. a fused feature vector) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of the input data item and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. Yang teaches the first inference model and the second inference model for a case where any of data items is unusable for the inference processing. One of ordinary skill would have motivation to combine Havaei, Garg and Yang to “improve[] in robustness of the multimodal machine learning system via training and using an odd-one-out network with a robust fusion layer” (Yang, Paragraph [0002]).
Regarding Claim 17,
The combination of Havaei, Garg and Yang teaches:
“A learning method of a neural network used for the control device according to claim 1, comprising:” (preamble)
“performing first learning using a plurality of learning data items which are a plurality of associated learning data items and are determined by the determination unit to be usable for the inference processing, and the first inference model which operates using the plurality of learning data items as inputs” (Havaei, Fig. 1,
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “In the context of HeMIS, the back end can be interpreted as learning to separately map each modality into an embedding common to all modalities, within which vector algebra operations carry well-defined semantics.”; Examiner’s note: performing first learning using a plurality of learning data items which are a plurality of associated learning data items (i.e. learning to map each modality into an embedding common to all modalities in the back end) and are determined by the determination unit to be usable for the inference processing (i.e. modalities available at inference time), and the first inference model (i.e. modality-specific convolutional neural network) which operates using the plurality of learning data items as an input (i.e. the plurality of modalities) is taught.)
“performing second learning using learning data out of a plurality of learning data items related to each other, which is determined by the determination unit to be usable for the inference processing, a second inference model which operates using learning data determined to be usable for the reference processing as an input, and the first inference model which performs an operation by a layer using the learning data determined to be usable for the reference processing as an input, and operates by inputting an output from the layer and an output from the second reference model to the connection layer (Garg, Col. 3, Lines 66-67 and Col. 4, Lines 1-5, Lines 15-24, Lines 38-44 and Lines 50-53, “the image classification model is trained in two phases. The first phase involves training an intermediate model as a multi-branch model with separate visual and text processing branches. Advantageously, the training in the first phase is done using a modality dropout approach in which the model is trained utilizing text data derived from some training images but not others […] a second phase may then be performed using only the visual processing branch. Advantageously, this second phase may be performed to refine the model for use as a model at inference. For example, the text-based processing branch is disabled or removed from the intermediate model to produce a model, and training continues using only image input (without OCR derivation of text) in order to fine tune the model and verify its accuracy before use of the generated image classification model at inference […] an intermediate model uses teacher student model distillation. In teacher student model distillation, the teacher model is typically a fairly complex and well-performing model, and the student model is typically a less complex model. For example, the teacher model may have more or wider hidden layers than the student model […] a teacher model is trained using a text-based processing branch and an image-based processing branch, and the student model is trained using only an image-based processing branch.”; Garg, Col. 13, Lines 28-32, “The result that is generated by the sets of convolutional layers and max pooling layers may be a matrix of numbers, such as floating-point numbers. The matrix may then be converted to a vector for processing by the set of fully-connected layers.”; Examiner’s note: performing second learning (i.e. second phase training) using learning data out of a plurality of learning data items (i.e. data out of first phase training with using a modality dropout approach) related to each other, which is determined by the determination unit (i.e. fine tuning the model) to be usable for the inference processing (i.e. use of generated image classification model at inference teaches second learning is usable for the inference processing), a second inference model (i.e. generated image classification model after fine tuning, the student model) which operates using learning data determined to be usable for the reference processing as an input (i.e. learning data after the first phase training and the disabled/removed text-based processing branch are usable for the reference processing as an input), and the first inference model (i.e. the teacher model) which performs an operation by a layer (i.e. the teacher model with hidden layers) using the learning data determined to be usable for the reference processing as an input (i.e. learning data from the first phase training to be used in the second phase), and operates by inputting an output from the layer (i.e. an output from the hidden layer) and an output from the second reference model (i.e. the student model) to the connection layer (i.e. the set of fully-connected layers) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 14 above and applicable herein.
Regarding Claim 18,
The combination of Havaei, Garg and Yang teaches:
“The method according to claim 17,” (preamble)
“wherein in the first learning, learning is performed by inputting a predetermined fixed value to the connection layer in place of the output from the second inference model” (Garg, Col. 9, Lines 17-20 and Lines 26-34, “In refining image classification model 160, the refinement may be based on performing additional training (e.g., one or more training epochs) and evaluation using a loss function (e.g., mean squared error loss) […] This process may be repeated in an iterative manner until a desired stopping point is reached. For example, the desired stopping point may correspond to satisfaction of an accuracy metric, exhaustion of a quantity of training time or training iterations, etc.”; Garg, Col. 10, Lines 52-55, “[…] the text processing branch is removed by changing the structure of the machine learning model itself to produce an image classification model with a single processing branch, rather than multiple processing branches”; Examiner’s note: wherein in the first learning (i.e. preforming additional training and evaluation using a loss function), learning is performed by inputting a predetermined fixed value (i.e. a desired stopping point corresponding to satisfaction of an accuracy metric, exhaustion of a quantity of training time or training iterations, etc.) to the connection layer in place of the output from the second inference model (i.e. an image classification model created by changing the structure of the machine learning model itself with a single processing branch, rather than multiple processing branches) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 17 above and applicable herein.
Regarding Claim 19,
The combination of Havaei, Garg and Yang teaches:
“The method according to claim 17,” (preamble)
“wherein in the second learning, learning is preformed by inputting a predetermined fixed value to the connection layer in place of an output from a layer of the plurality of layers, which uses learning data determined by the determination unit to be unusable for the inference processing” (Yang, Fig. 7,
PNG
media_image3.png
401
489
media_image3.png
Greyscale
; Yang, Paragraphs [0041], [0049] and [0050], “After the machine-learning algorithm 310 achieves a predetermined performance level (e.g., 100% agreement with the outcomes associated with the training dataset 312), the machine-learning algorithm 310 may be executed using data that is not in the training dataset 312. The trained machine-learning algorithm 310 may be applied to new datasets to generate annotated data […] The multimodal fusion network 602 receives input modalities 604a, 604b, 604c and extracts features 606a, 606b, 606c from each modality that are feature vectors. The output of the feature extractors 606 is fed into an odd-one-out network 612. The odd-one-out network 612 generates an “inconsistent” modality prediction that is fed to a robust fusion layer 608 along with the output of the feature extractors 606. The robust fusion layer 608 outputs a fused feature vector that is subsequently fed to downstream layers 610 to produce an output […] The network 700 receives features 702 such as output from feature extractor 602a, 602b, and 602c and generates a modality prediction weights 704 such that for each feature channel is an associated modality prediction weight 704a, 704b, and 704c. These modality prediction weights 704a, 704b, and 704c produce an odd-one-out vector that is forwarded to the robust feature fusion layer.”; Examiner’s note: wherein in the second learning (i.e. training the machine learning using data that is not in the training dataset), learning is performed by inputting a predetermined fixed value (i.e. a predetermined performance level (e.g., 100% agreement with the outcomes associated with the training dataset)) to the connection layer (i.e. feature extractor layer) in place of an output from a layer of the plurality of layers (i.e. generated annotated data), which uses learning data (i.e. data not meeting the predetermined performance level) determined by the determination unit (i.e. odd-one-out network) to be unusable for the inference processing (i.e. odd-one-out vector) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 17 above and applicable herein.
Regarding Claim 20,
Havaei teaches:
controlling an operation of a first inference model that is formed by a plurality of layers to which the plurality of data items are input, respectively, for extraction of feature values of input data items, and a connection layer which outputs output data as an inference result based on the extracted feature values (Havaei, Fig. 1,
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others.”; Examiner’s note: controlling an operation of a first inference model (i.e. modality-specific convolutional neural network) that is formed by a plurality of layers to which the plurality of data items are input (i.e. modality-specific convolutional layers), respectively, for extraction of feature values of input data items (i.e. extraction of feature maps), and a connection layer (i.e. convolutional layer in front end) which outputs output data as an inference result (i.e. classifications outputs) based on the extracted feature values (i.e. extracted feature maps) is taught.)
“determining whether or not each of the plurality of data items is usable for the inference processing” (Havaei, Fig. 1 and Section 1, “This approach presents the advantage of being robust to any combinatorial subset of available modalities provided as input, without the need to learn a combinatorial number of imputation models.”; Havaei, Section 2, “Here, we make the HeMIS architecture robust to missing modalities by randomly dropping any number for a given training example […] we start randomly dropping modalities, ensuring a higher probability of dropping zero or one modality only.”; Examiner’s note: determining whether or not each of the plurality of data items is usable for the inference processing (i.e. randomly dropping modalities for missing modalities and any combinatorial subset of available modalities provided as input teach input is usable for the inference processing) is taught. See Fig. 1 in the above limitation of claim 20.)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the first inference model to output the output data from the connection layer based on a plurality of feature values extracted by the plurality of layers (Havaei, Fig. 1 and Section 2, “We propose an approach wherein each modality is initially processed by its own convolutional pipeline, independently of all others. After a few independent stages, feature maps from all available modalities are merged by computing map-wise statistics such as the mean and the variance, quantities whose expectation does not depend on the number of terms (i.e. modalities) that are provided. After merging, the mean and variance feature maps are concatenated and fed into a final set of convolutional stages to obtain network output.”; Examiner’s note: controlling (i.e. convolutional pipeline), in a case where it is determined by said determining (i.e. the convolutional neural network) that all of the plurality of data items are usable for the inference processing (i.e. all available modalities used for the inference processing), the first inference model (i.e. modality-specific convolutional neural network) to output the output data from the connection layer (i.e. output obtained at a final set of convolutional stages) based on a plurality of feature values extracted by the plurality of layers (i.e. extracted feature maps from all available modalities in convolutional layers) is taught. See Fig. 1 in the above limitation of claim 20.)
Havaei does not explicitly teach:
“A non-transitory computer-readable storage medium storing a program for causing a computer to execute a method of controlling a control device that performs inference processing using a plurality of data items related to each other as inputs, wherein the method comprises:”
controlling an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values
“controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input”
Garg teaches:
“A non-transitory computer-readable storage medium storing a program for causing a computer to execute a method of controlling a control device that performs inference processing using a plurality of data items related to each other as inputs, wherein the method comprises:” (Garg, Col. 16, Lines 54-59, “Each such computing device typically includes a processor (or multiple processors) that executes program instructions or modules stored in a memory or other non-transitory computer-readable storage medium or device (e.g., solid state storage devices, disk drives, etc.)”)
controlling an operation of a second inference model using any of the plurality of data items as an input for extraction of a feature value of the input data item (Garg, Col. 11, Lines 40-46, “An image set 320 also may be used in the student model 354 as part of teacher student training system 350. Image set 320 may constitute a training set of all of the images used in training the models within teacher student training system 350. Image set 320 may also constitute a subset of images in a larger training set. Images in the training set may be obtained from users.”; Garg, Col. 13, Lines 23-25 and Lines 28-32, “[…] the student model 354 can be trained by processing image set 320 using a visual processor 322 […] The result that is generated by the sets of convolutional layers and max pooling layers may be a matrix of numbers, such as floating-point numbers. The matrix may then be converted to a vector for processing by the set of fully-connected layers.”; Examiner’s note: controlling an operation of a second inference model (i.e. the student model) using any of the plurality of data items as an input (i.e. images obtained from users) for extraction of a feature value of the input data item (i.e. a vector for processing by the set of fully-connected layers)
controlling, in a case where it is determined by said determining that all of the plurality of data items are usable for the inference processing, the second inference model not to perform extraction of feature values (Garg, Col. 3, Lines 21-23, Lines 37-40 and Lines 46-49, “[…] the image set generation system may be configured to determine whether all of the image classes in the preferred image set for the item are present […] the image classification model may be implemented as a multi-branch model that uses different processing branches to process input data in different modalities […] a multi-branch model may include an image-based processing branch for input in an image modality, and a text-based processing branch for input in a text-based modality.”; Garg, Col. 10, Lines 44-46, “If the desired stopping point has been reached at block 208, then at block 210 the text processing branch can be removed from the image classification model.”; Examiner’s note: controlling (i.e. an image-based processing branch and a text-based processing branch), in a case where it is determined by said determining (i.e. the image set generation system) that all of the plurality of data items are usable for the inference processing (i.e. all of the image classes in the preferred image set for the item are present), the second inference model not to perform extraction of feature values (i.e. if the desired stopping point reached, the text processing branch is removed from the image classification model which means the image classification model not performing the text processing branch) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of the input data item and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. One of ordinary skill would have motivation to combine Havaei and Garg so that “the speed of the system may in some implementations be improved while maintaining an accuracy level” (Garg, Col. 4, Lines 24-26).
Yang teaches:
“controlling, in a case where it is determined by said determining that any of the plurality of data items is unusable for the inference processing, the first inference model to output the output data from the connection layer, based on a feature value extracted by each layer to which data determined by the determination unit to be usable for the inference processing is input, and a feature value extracted by the second inference model using the data determined by said determining to be usable for the inference processing as an input” (Yang, Paragraphs [0049] and [0050], “The multimodal fusion network 602 receives input modalities 604a, 604b, 604c and extracts features 606a, 606b, 606c from each modality that are feature vectors. The output of the feature extractors 606 is fed into an odd-one-out network 612. The odd-one-out network 612 generates an “inconsistent” modality prediction that is fed to a robust fusion layer 608 along with the output of the feature extractors 606. The robust fusion layer 608 outputs a fused feature vector that is subsequently fed to downstream layers 610 to produce an output […] The network 700 receives features 702 such as output from feature extractor 602a, 602b, and 602c and generates a modality prediction weights 704 such that for each feature channel is an associated modality prediction weight 704a, 704b, and 704c. These modality prediction weights 704a, 704b, and 704c produce an odd-one-out vector that is forwarded to the robust feature fusion layer.”; Yang, Fig. 6,
PNG
media_image2.png
480
656
media_image2.png
Greyscale
; Examiner’s note: controlling (i.e. the multimodal fusion network), in a case where it is determined by said determining (i.e. the odd-one-out network) that any of the plurality of data items is unusable for the inference processing (i.e. an odd-one-out vector teaches a data item is unusable for the inference processing), the first inference model to output the output data (i.e. the output of the feature extractors 606) from the connection layer (i.e. feature extractor layer), based on a feature value extracted by each layer (i.e. extracted feature 606a, 606b, 606c) to which data determined by the determination unit (i.e. the feature extractor) to be usable for the inference processing is input (i.e. feature vectors), and a feature value extracted by the second inference model (i.e. a fused feature vector by the robust fusion layer) using the data determined by said determining (i.e. the odd-one-out network) to be usable for the inference processing as an input (i.e. a fused feature vector) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of the input data item and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. Yang teaches the first inference model and the second inference model for a case where any of data items is unusable for the inference processing. One of ordinary skill would have motivation to combine Havaei, Garg and Yang to “improve[] in robustness of the multimodal machine learning system via training and using an odd-one-out network with a robust fusion layer” (Yang, Paragraph [0002]).
Claims 6 and 8-12 are rejected under 35 U.S.C. 103 as being unpatentable over Havaei, in view of Garg, and further in view of Yang as applied in claim 1, and further in view of Sekiguchi et al. (EP 3439282 A1) (hereinafter Sekiguchi).
Regarding Claim 6,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 5,” (preamble)
The combination of Havaei, Garg and Yang does not explicitly teach:
“wherein the evaluation value is a value related to one of a luminance of each image data item, a blur amount of the image data item, and a high-sensitivity noise amount of the image data item” (preamble)
Sekiguchi teaches:
“wherein the evaluation value is a value related to one of a luminance of each image data item, a blur amount of the image data item, and a high-sensitivity noise amount of the image data item” (Sekiguchi, Paragraphs [0023] and [0295], “These image capture conditions include the exposure conditions described above (i.e. the charge accumulation time, the gain, the ISO sensitivity, the frame rate, and so on) and the image processing conditions described above (for example, a parameter for white balance adjustment, a gamma correction curve, a parameter for display luminance adjustment, a saturation adjustment parameter, and so on) […] the generation unit 33c of the image processing unit 33 compensates for blurring of the boundary between elements of the photographic subject described above by performing contrast adjustment processing in addition to the noise reduction processing, or along with the noise reduction processing.”; Examiner’s note: wherein the evaluation value (i.e. image processing conditions) is a value related to one of a luminance of each image data item (i.e. a parameter for display luminance adjustment), a blur amount of the image data item (i.e. blurring of the boundary between elements of the photographic subject), and a high-sensitivity noise amount of the image data item (i.e. the noise reduction processing) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item, a control unit to control operations of the first inference model and the second inference model and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. Yang teaches the first inference model and the second inference model for a case where any of data items is unusable for the inference processing. Sekiguchi teaches the evaluation value is a value related to one of a luminance of each image data item, a blur amount of the image data item, and a high-sensitivity noise amount of the image data item. One of ordinary skill would have motivation to combine Havaei, Garg Yang and Sekiguchi to “suppress degradation of the accuracy of detection of the elements of the photographic subject due to difference between the image capture conditions for the various blocks” (Sekiguchi, Paragraph [0268]).
Regarding Claim 8,
The combination of Havaei, Garg, Yang and Sekiguchi teaches:
“The control device according to claim 4,” (preamble)
“wherein the plurality of data items are a plurality of image data items obtained by capturing an image of a person at different angles” (Sekiguchi, Paragraph [0451], “The lens movement control unit 34d adjusts the angle of view by the image capture optical system 31 by shifting the zoom lens in the direction of the optical axis. In other words, by shifting the zoom lens, it is possible to perform adjustment of the image produced by the image capture optical system 31 so as to obtain an image of the photographic subject over a wide range, to obtain a large image for a faraway photographic subject, and the like.”; Examiner’s note: wherein the plurality of data items are a plurality of image data items (i.e. an image of the photographic subject over a wide range) obtained by capturing an image of a person at different angles (i.e. obtaining an image of the subject by adjusting the angle of view) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 4 above and applicable herein.
Regarding Claim 9,
The combination of Havaei, Garg, Yang and Sekiguchi teaches:
“The control device according to claim 8,” (preamble)
“wherein the evaluation value is a value related to an orientation of a face of a person whose image appears in the image data” (Sekiguchi, Paragraphs [0022] and [0030], “By performing per se known object recognition processing, the object detection unit 34a detects elements of the photographic subject from the image acquired by the image capture unit 32, such as a person (i.e. the face of a person) […] In some of the subsequent figures coordinate axes are displayed so that, taking the coordinate axes shown in Fig. 2 as reference, the orientation of each figure can be understood.”; Examiner’s note: wherein the evaluation value (i.e. performing object recognition processing) is a value related to an orientation of a face of a person (i.e. the orientation of the face of a person acquired by the image capture unit) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein.
Regarding Claim 10,
The combination of Havaei, Garg, Yang and Sekiguchi teaches:
“The control device according to claim 4,” (preamble)
“wherein the plurality of data items include image data of interest and reference image data for performing noise reduction for eliminating noise from the image data of interest” (Sekiguchi, Paragraph [0115], “The predetermined image processing is processing for calculating the main image data item of the position for attention in the image that is the subject of processing by referring to the main image data items in a plurality of reference positions around the position for attention, and may include, for example, pixel defect correction processing, color interpolation processing, contour enhancement processing, noise reduction processing, and so on.”; Examiner’s note: wherein the plurality of data items (i.e. the main image data items) include image data of interest (i.e. the main image data item of the position for attention in the image) and reference image data for performing noise reduction for eliminating noise from the image data of interest (i.e. referring to the main image data items in a plurality of reference positions around the position for attention including noise reduction processing) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 4 above and applicable herein.
Regarding Claim 11,
The combination of Havaei, Garg, Yang and Sekiguchi teaches:
“The control device according to claim 10,” (preamble)
“wherein the evaluation value is a value related to a difference between the image data of interest and the reference image data” (Sekiguchi, Paragraphs [0161] and [0380], “The lens movement control unit 34d of the control unit 34 performs focus detection processing by employing the signal data (i.e. the image data) corresponding to a predetermined position upon the imaging screen (i.e. the point of focusing) […] It should be understood that if the second correction processing is performed upon that signal data that, among the signal data, was captured under the second image capture conditions in order to reduce the difference between the signal data after the second correction processing and the signal data that was captured under the first image capture conditions […]”; Examiner’s note: wherein the evaluation value (i.e. performing the second correction processing) is a value related to a difference between the image data of interest and the reference image data (i.e. different between the signal data after the second correction processing (i.e. the image data of interest) and the signal data captured under the first image capture conditions (i.e. the reference image data) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 10 above and applicable herein.
Regarding Claim 12,
The combination of Havaei, Garg, Yang and Sekiguchi teaches:
“The control device according to claim 1,” (preamble)
“wherein the plurality of data items are image data items obtained by capturing an image of an object” (Sekiguchi, Paragraph [0022], “By performing per se known object recognition processing, the object detection unit 34a detects elements of the photographic subject from the image acquired by the image capture unit 32, such as a person (i.e. the face of a person)”; Examiner’s note: wherein the plurality of data items are image data items obtained by capturing an image of an object (i.e. elements of the photographic subject from the image acquired by the image capture unit) is taught.)
“wherein the determination unit determines whether or not each image data item is usable for the inference processing, based on image-capturing conditions at the time of capturing the image data items” (Havaei, Fig. 1 and Section 1, “This approach presents the advantage of being robust to any combinatorial subset of available modalities provided as input, without the need to learn a combinatorial number of imputation models.”;
PNG
media_image1.png
326
447
media_image1.png
Greyscale
; Havaei, Section 2, “Here, we make the HeMIS architecture robust to missing modalities by randomly dropping any number for a given training example […] we start randomly dropping modalities, ensuring a higher probability of dropping zero or one modality only.”; Sekiguchi, Paragraphs [0023], “These image capture conditions include the exposure conditions described above (i.e. the charge accumulation time, the gain, the ISO sensitivity, the frame rate, and so on) and the image processing conditions described above (for example, a parameter for white balance adjustment, a gamma correction curve, a parameter for display luminance adjustment, a saturation adjustment parameter, and so on).”; Examiner’s note: wherein the determination unit (i.e. abstraction layer) determines whether or not each image data item is usable for the inference processing (i.e. randomly dropping modalities for missing modalities and any combinatorial subset of available modalities provided as input teach input is usable for the inference processing), based on image-capturing conditions at the time of capturing the image data items (i.e. image capture conditions including the exposure conditions such as the charge accumulation time, the gain, the ISO sensitivity, the frame rate, and so on) is taught.)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Havaei, in view of Garg, and further in view of Yang as applied in claim 1, and further in view of Minoru. (JP 6061782 B2).
Regarding Claim 7,
The combination of Havaei, Garg and Yang teaches:
“The control device according to claim 5,” (preamble)
The combination of Havaei, Garg and Yang does not explicitly teach:
“wherein the evaluation value is a value related to a noise amount of a feature component of the sound data”
Minoru teaches:
“wherein the evaluation value is a value related to a noise amount of a feature component of the sound data” (Minoru, Paragraphs [0013] and [0038], “normalization means for normalizing an evaluation value by the one-class support vector machine (e.g., different in FIG. 6 a sound determination unit 230) […] an indication of the abnormal sound discrimination performance becomes a value of 1 or more if the length of the longest branch of feature vectors of abnormal noise, also larger the value, the clearly abnormal sound is determined show.”; Examiner’s note: wherein the evaluation value (i.e. an evaluation value) is a value related to a noise amount of a feature component of the sound data (i.e. the length of the longest branch of feature vectors of abnormal noise) is taught.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the hetero-modal image segmentation in Havaei, and the image classification with modality dropout as taught in Garg. Havaei teaches a first inference model formed by a plurality of layers to which the plurality of data items are input and a connection layer which outputs output data. Garg teaches a second inference model to which any of the plurality of data items are input for extraction of a feature value of each of the any input data item, a control unit to control operations of the first inference model and the second inference model and the second inference model not to perform extraction of feature values when all data items are usable for the inference processing. Yang teaches the first inference model and the second inference model for a case where any of data items is unusable for the inference processing. Minoru teaches the evaluation value related to a noise amount of a feature component of the sound data. One of ordinary skill would have motivation to combine Havaei, Garg Yang and Minoru to “improve the accuracy of the abnormal sound detection” (Minoru, Paragraph [0061]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Neverova et al. teaches adaptive multi-modal gesture recognition. Tran et al. teaches missing modalities imputation via cascaded residual autoencoder.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YONG D RHO whose telephone number is (571)270-0194. The examiner can normally be reached 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 5712705871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YONG DOO RHO/Examiner, Art Unit 2147 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148