Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “an inference unit that executes inference…”, “a determination unit that determines…”, and “a decision unit that decides…” in claim 1, “the decision unit calculates… and determines…” in claim 2, “the decision unit outputs…” in claim 3, “the decision unit changes a result…” in claim 4, “a calculation unit that calculates…” and “the decision unit decides the substitute data…” in claims 5 and 6, “the decision unit decides… based on characteristic information…” in claim 10, “the decision unit determines… by comparing…” in claim 14, “the decision unit obtains a difference…” in claim 15, and “the inference unit executes… a trained model and…” in claim 16.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 is/are rejected under 35 U.S.C. 101 because they are directed to an abstract idea without significantly more.
Regarding claims 1-20,
Step 1: The preamble of claims 1-18 recites an apparatus, which falls within the statutory category of a system. The preamble of claim 19 recites a method, which falls within the statutory category of a process. The preamble of claim 20 recites a non-transitory computer-readable storage medium, which falls within the statutory category of a manufacture.
Regarding claim 1,
Step 2A – Prong One: Claim 1 recites:
An inference apparatus comprising:
an inference unit that executes inference processing using a plurality of pieces of inference data;
a determination unit that determines, for each of a plurality of pieces of input data inputted into the inference apparatus, whether each of the plurality of pieces of input data are suitable for the inference processing or unsuitable for the inference processing; and
a decision unit that decides, in place of input data determined to be unsuitable, substitute data based on input data determined to be suitable,
wherein the inference unit applies the input data determined to be suitable and the substitute data to an inference model as the plurality of pieces of inference data and executes the inference processing.
The broadest reasonable interpretation, in light of the Specification, of the bolded limitations above are mental processes that can be performed in the human mind with or without a pen and/or paper. A human could use observation, evaluation, and judgement to determine whether each of the plurality of pieces of data input into an inference apparatus are suitable or unsuitable for inference processing. A human could use evaluation and judgement to decide substitute data based on suitable input data to substitute unsuitable input data.
Step 2A – Prong One (Yes).
Step 2A – Prong Two: The additional element of the claim regarding “an inference unit that executes inference processing using a plurality of pieces of inference data;” and “wherein the inference unit applies the input data determined to be suitable and the substitute data to an inference model as the plurality of pieces of inference data and executes the inference processing” are mere instructions to apply the judicial exception on a generic computer (See MPEP 2106.05(f)). Even when viewed in combination, the additional element fails to integrate the judicial exception into a practical application.
Step 2A – Prong Two (No).
Step 2B: The additional element of the claim regarding “an inference unit that executes inference processing using a plurality of pieces of inference data;” and “wherein the inference unit applies the input data determined to be suitable and the substitute data to an inference model as the plurality of pieces of inference data and executes the inference processing” are mere instructions to apply the judicial exception on a generic computer (See MPEP 2106.05(f)). The computer is recited at a high level of generality and imposes no meaningful limitations on the claim. Even when viewed in combination, the additional element fails to amount to significantly more than the judicial exception.
Step 2B (No).
Claim 1 is ineligible.
Regarding claims 19-20,
These claims are similar in scope to claim 1 and are rejected under a similar rationale. The processors recited in these claims are also generic computing components.
Claims 19-20 are ineligible.
Dependent claims:
Claim 2: This claim recites further abstract ideas (mental processes, “determines… the suitability for when that piece of input data is used…”). The limitation regarding “the decision unit calculates… evaluation information…” is a mathematical calculation and thus falls within the mathematical concepts grouping of abstract ideas (See MPEP 2106.04(a)(2)(I)(A)). There are no additional elements in the claim that integrate the judicial exception into a practical application and there are no additional elements in the claim that would amount to significantly more than the judicial exception. Thus, this claim is ineligible.
Claims 3, 7-13, and 17: These claims recite further additional elements that are insignificant extra-solution activities that amount to mere data gathering (See MPEP 2106.05(g)). Data gathering is a well-understood, routine conventional activity as recognized by the courts (See MPEP 2106.05(d)(II)). Thus, these claims fail to integrate the judicial exception into a practical application and fail to amount to significantly more than the judicial exception. Thus, these claims are ineligible.
Claims 4-6 and 14-15: These claims only recite further abstract ideas (mental processes, mathematical concepts) and thus are ineligible.
Claims 16 and 18: These claims recite further instructions to apply the judicial exception on a generic computer. The computer is recited at a high level of generality and imposes no meaningful limitations on these claims. These claims fail to integrate the judicial exception into a practical application and fail to amount to significantly more than the judicial exception. Thus, these claims are ineligible.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1 and 16-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhao et al. (NPL: Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities, published Aug. 2021, hereinafter “Zhao”).
Regarding claim 1, Zhao teaches an inference apparatus (Zhao, Section 4.5) comprising:
an inference unit that executes inference processing using a plurality of pieces of inference data (Zhao, Section 3 Paragraph 1 – “Our proposed method aims to recognize the emotion category yi for every video segment si with full modalities, or with only partial modalities available, for the example shown in Figure 1, there exist only acoustic and textual modalities when visual modality is missing.” and in Fig. 2 – teaches an inference unit that executes inference processing using a plurality of pieces of inference data (Fig. 2 shows model with classifier, and thus an inference unit, that executes inference processing using a plurality of pieces of inference data, the plurality of data in Fig. 1 being xa, xv and xt));
a determination unit that determines, for each of a plurality of pieces of input data inputted into the inference apparatus, whether each of the plurality of pieces of input data are suitable for the inference processing or unsuitable for the inference processing (Zhao, Fig. 2 – teaches determining, for each of a plurality of pieces of input data input into the inference apparatus, whether each of the plurality of pieces of input data are suitable for inference processing or unsuitable for inference processing (Fig. 2 shows determination of data that is either missing (such as xvmiss indicating visual data is missing) or available (such as xa and xt indicating that acoustic and textual data are available), and thus determines if each of the plurality of pieces of input data are suitable or unsuitable)); and
a decision unit that decides, in place of input data determined to be unsuitable, substitute data based on input data determined to be suitable (Zhao, Section 3.1.3 Paragraph 1 – “We propose an autoencoder-based Imagination Module to predict the multimodal embeddings of the missing modalities given the multimodal embeddings of the available modalities.” – teaches a decision unit that decides, in place of input data determined to be unsuitable, substitute data based on suitable input data (predicts embeddings of missing modalities based on embeddings of available modalities, thus deciding substitute data based on suitable data in place of unsuitable data)),
wherein the inference unit applies the input data determined to be suitable and the substitute data to an inference model as the plurality of pieces of inference data and executes the inference processing (Zhao, Fig. 2 (c) and Fig. 2 description – “(c) Missing Modality Imagination Network (MMIN) at the inference stage (taking the visual modality missing condition as an example). MMIN can inference under different missing modality conditions.” – teaches wherein the inference unit applies the input data determined to be suitable and the substitute data (substitute data as in Zhao at Section 3.1.3 Paragraph 1) to an inference model as the plurality of pieces of inference data and executes the inference processing (Fig. 2 (c) shows execution of inference processing using the suitable input data and substitute data)).
Claims 19-20 incorporate substantively all the limitations of claim 1 in a method and non-transitory computer-readable storage medium and are rejected on similar grounds as above. Zhao teaches the processors of these claims at Section 4.5.
Regarding claim 16, Zhao teaches the apparatus according to claim 1, wherein
the inference unit executes the inference processing in which a trained inference model and parameters are used (Zhao, Fig. 2 (c) and Fig. 2 description – “(c) Missing Modality Imagination Network (MMIN) at the inference stage (taking the visual modality missing condition as an example). MMIN can inference under different missing modality conditions.” – teaches wherein the inference unit executes the inference processing in which a trained inference model and parameters are used (Fig. 2 (c) shows the inference stage in which a trained inference model and parameters are used, and Zhao at Section 4.5 further states that all models are run on GPU)), and
regarding the inference model, training processing is executed in advance using training data including training input data and supervisory data (Zhao, Fig. 2 (a) and Fig. 2 description – “(a) MMIN at the training stage (taking the visual modality missing condition as example). MMIN is trained with all six possible missing modality conditions (Table 1).” – teaches wherein training processing of the inference model is executed in advance using training data (model trained with all six possible missing modality conditions) and supervisory data (Zhao teaches supervisory, or target, data at Section 3 Paragraph 1 – “We denote the target set Y…”)).
Regarding claim 17, Zhao teaches the apparatus according to claim 16, wherein
the training input data includes a data set for which a plurality of pieces of input data determined to be suitable by the decision unit have been combined or a data set for which input data determined to be suitable by the decision unit and substitute data decided based on characteristic information of that input data have been combined (Zhao, Section 4.1.1 Paragraph 1 – “We first define the original training set which contains all the three modalities as the full-modality training set.” – teaches wherein the training input data includes a dataset for which a plurality of pieces of input data determined to be suitable by the decision unit have been combined (original training data set contains all three modalities, and thus is determined to be suitable, as the full-modality training set, thus teaching wherein the training input data includes a data set for which a plurality of suitable inputs are combined)).
Regarding claim 18, Zhao teaches the apparatus according to claim 1, wherein
the inference model is a neural network, and the inference processing is deep learning using a neural network (Zhao, Fig. 2, Fig. 2 description – “Illustration of the Missing Modality Imagination Network (MMIN) framework.”, and Section 4.5 Paragraph 1 – “All models are implemented with Pytorch deep learning toolkit…” – teaches wherein the inference model is a neural network, and the inference processing is deep learning using a neural network (all models implemented using deep learning, thus inference processing is deep learning using a neural network)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2-3, 5-10, and 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao in view of Leszcuk et al. (NPL: Key Indicators for Monitoring of Audiovisual Quality, published 2014, hereinafter “Leszcuk”).
Regarding claim 2, Zhao teaches the apparatus according to claim 1.
Zhao fails to explicitly teach wherein the decision unit calculates, for each of the plurality of pieces of input data, evaluation information for that input data, and determines, for each of the plurality of pieces of input data, the suitability for when that piece of input data is used in the inference processing by comparing the evaluation information with a predetermined threshold.
However, analogous to the field of the claimed invention, Leszczuk teaches: wherein
the decision unit calculates, for each of the plurality of pieces of input data, evaluation information for that input data (Leszczuk, Fig. 1, Section IV Subsection A & C – teaches wherein the decision unit calculates, for each of the plurality of pieces of input data, evaluation information for that input data (Fig. 1 shows a unit that calculates metrics for each of the plurality of pieces of input data, Section IV Subsection A & C describe video and audio artifact metrics such as metrics for blurring and audio muting)), and
determines, for each of the plurality of pieces of input data, the suitability for when that piece of input data is used in the inference processing by comparing the evaluation information with a predetermined threshold (Leszczuk, Fig. 1, Section IV Subsection A & C – teaches determining for each of the plurality of pieces of input data, the suitability for when that piece of input data is used in inference processing by comparing the evaluation information with a predetermined threshold (Fig. 1 shows comparing the metrics for the plurality of pieces of input data to predetermined thresholds to determine if input meets quality standards, Section IV Subsections A & C describe thresholds for metrics such as exposure time distortion and muting)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the calculations of evaluation information and predetermined thresholds of Leszczuk to the inference model and decision unit of Zhao. Doing so would provide automatic quality measurements associated with occurrences of perceptible degradation in audio/video (Leszczuk, Introduction)
Regarding claim 3, the combination of Zhao and Leszczuk teaches the apparatus according to claim 2, wherein
in a case where all of the plurality of pieces of input data are determined to be suitable by the decision unit, the decision unit outputs those pieces of input data determined to be suitable as the plurality of pieces of inference data to the inference unit (Zhao, Section 3 Paragraph 1 – “Our proposed method aims to recognize the emotion category yi for every video segment si with full modalities…” and Fig. 2 – teaches wherein in a case where all the plurality of pieces of input data are determined to be suitable by the decision unit, the decision unit outputs those pieces of input data determined to be suitable as the plurality of pieces of inference data to the inference unit (aims to recognize emotion category, and thus execute inference, for every segment with full modalities available. Thus, teaching in a case where all pieces of input data are suitable, or full modalities are available, outputting those pieces of input data as inference data to the inference unit)).
Regarding claim 5, the combination of Zhao and Leszczuk teaches the apparatus according to claim 2, further comprising:
wherein the decision unit decides the substitute data based on characteristic information of a respective piece of input data (Zhao, Section 3.1.3 Paragraph 1 – “We propose an autoencoder-based Imagination Module to predict the multimodal embeddings of the missing modalities given the multimodal embeddings of the available modalities.” – teaches a decision unit that decides substitute data based on characteristic information of suitable input data (predicts missing modalities based on available modalities, thus deciding substitute data based on characteristics of available modalities)).
Zhao fails to explicitly teach a calculation unit that calculates, for each of the plurality of pieces of input data, characteristic information of that piece of input data.
However, analogous to the field of the claimed invention, Leszczuk teaches:
a calculation unit that calculates, for each of the plurality of pieces of input data, characteristic information of that piece of input data (Leszczuk, Fig. 1 and Section IV Subsections A & C – teaches a calculation unit (metrics unit of MOAVI in Fig. 1) that calculates, for each of the plurality of pieces of input data, characteristic information of that piece of input data (calculates video and audio metrics such as blurring, exposure time distortion, clipping, and muting metrics)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the characteristic information calculations of Leszczuk to the decision unit and substitute data of Zhao in order to decide substitute data on characteristic information. Doing so would take the nature and characteristics of the video and audio artifacts into consideration for quality evaluations (Leszczuk Section IV Paragraph 1).
Regarding claim 6, the combination of Zhao and Leszczuk teaches the apparatus according to claim 5, wherein
the decision unit decides the substitute data based on a plurality of pieces of characteristic information of a respective piece of input data (Zhao, Section 3.1.3 Paragraph 1 – “We propose an autoencoder-based Imagination Module to predict the multimodal embeddings of the missing modalities given the multimodal embeddings of the available modalities.” – teaches a decision unit that decides the substitute data based on pieces of characteristic information of suitable input data (predicts missing modalities based on available modalities, thus deciding substitute data based on characteristics of the available modalities)).
Zhao fails to explicitly teach the calculation unit calculates, for each of the plurality of pieces of input data, a plurality of pieces of characteristic information of that input data.
However, analogous to the field of the claimed invention, Leszczuk teaches:
the calculation unit calculates, for each of the plurality of pieces of input data, a plurality of pieces of characteristic information of that input data (Leszczuk, Fig. 1 and Section IV Subsections A & C – teaches a calculation unit (metrics unit of MOAVI in Fig. 1) that calculates, for each of the plurality of pieces of input data, a plurality of pieces of characteristic information of that input data (calculates video and audio metrics such as blurring, exposure time distortion, clipping, and muting metrics, and thus calculates a plurality of pieces of characteristic information)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the plurality of pieces of characteristic information calculations of Leszczuk to the decision unit and substitute data of Zhao in order to decide substitute data on a plurality of pieces characteristic information. Doing so would take the nature and characteristics of the video and audio artifacts into consideration for quality evaluations (Leszczuk Section IV Paragraph 1).
Regarding claim 7, the combination of Zhao and Leszczuk teaches the apparatus according to claim 5, wherein
the characteristic information is a value for which a characteristic obtained for a respective piece of input data has been normalized by the calculation unit (Leszczuk, Section IV Subsections A & C – teaches wherein the characteristic information is a value (such as blurring, exposure time distortions, clipping, and muting metrics) for which a characteristic obtained for a respective piece of input data has been normalized by the calculation unit (teaches calculating artifact metrics such as blurring, ringing, exposure time distortions, clipping, and muting with thresholds, thus the characteristic information is a value, or metric, for which a characteristic obtained for a respective piece of input data has been normalized, consistent with the Specification at [0083])).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the characteristic information calculations of Leszczuk to the data and apparatus of Zhao. Doing so would take the nature and characteristics of the video and audio artifacts into consideration for quality evaluations (Leszczuk Section IV Paragraph 1).
Regarding claim 8, the combination of Zhao and Leszczuk teaches the apparatus according to claim 5, wherein
the plurality of pieces of input data include image data and sound data (Zhao, Section 3 Paragraph 1 – “Given a set of video segments S, we use x= (xa, xv, xt) to represent the raw multimodal features for a video segment s∈S, where xa, xv and xt represent the raw features of acoustic, visual and textual modalities respectively.” – teaches wherein the plurality of pieces of input data include image and sound data (data includes visual and acoustic modalities)).
Regarding claim 9, the combination of Zhao and Leszczuk teaches the apparatus according to claim 8, wherein
the evaluation information and the characteristic information include, in a case where the input data is the image data, a value related to either a blur amount, a luminance, or noise according to sensitivity at the time of capturing the image data (Leszczuk, Section IV Subsection A Paragraph 2 – “Blurring shows as reduced sharpness of edges and spatial detail. In compressed video, it results from a loss of high frequency information during coding. Measurement of this artifact is based on the cosine of the angle between perpendiculars to planes in adjacent pixels which is a good characteristic of picture smoothness [17].” – teaches wherein the evaluation information and the characteristics information include, in a case where the input data is the image data, a value related to either a blur amount, a luminance, or noise according to sensitivity at the time of capturing the image data (calculates blurring metrics based on cosine of angle between perpendiculars to planes in adjacent pixels)), and
in a case where the input data is the sound data, a value related to a sound volume or noise sound (Leszczuk, Section IV Subsection C Paragraph 2 – “Mute– signal losses are one of the most common degradation in audio streaming at low bit rates. The detection algorithm is based on two different thresholds, once the signal has frequency components only in the hearing range (between 20 Hz and 20 kHz). These thresholds are: the minimum level of signal noticeable by the human ear, and the duration of the shortest silent interval perceptible as a differentiated mute. The value of these thresholds must be adaptive to the nature of the sound, as experimental test have demonstrated.” – teaches the evaluation information and the characteristic information include, in a case where the input data is the sound data, a value related to a sound volume or a noise sound (detects muting based on two different thresholds, where the thresholds are the minimum level of signal and duration of shortest silent interval, thus determining a value related to sound volume)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the characteristic information of Leszczuk to the data and apparatus of Zhao. Doing so would take the nature and characteristics of the video and audio artifacts into consideration for quality evaluations (Leszczuk Section IV Paragraph 1).
Regarding claim 10, the combination of Zhao and Leszczuk teaches the apparatus according to claim 9, wherein
the decision unit decides the substitute data based on the first characteristic information and second characteristic information of a respective piece of input data (Zhao, Section 3.1.3 Paragraph 1 – “We propose an autoencoder-based Imagination Module to predict the multimodal embeddings of the missing modalities given the multimodal embeddings of the available modalities.” – teaches a decision unit that decides the substitute data based on characteristics of input data (predicts missing modalities based on available modalities, thus deciding substitute data based on characteristics of the available modalities)).
Zhao fails to explicitly teach the characteristic information includes, in a case where the input data is the image data, first characteristic information related to an amount of noise according to sensitivity at the time of capturing the image data and second characteristic information related to a blur amount of the image data, and in a case where the input data is the sound data, first characteristic information related to noise sound and second characteristic information related to a sound volume.
However, analogous to the field of the claimed invention, Leszczuk teaches:
the characteristic information includes, in a case where the input data is the image data, first characteristic information related to an amount of noise according to sensitivity at the time of capturing the image data and second characteristic information related to a blur amount of the image data (Leszczuk, Section IV Subsection A Paragraph 2 – “Blurring shows as reduced sharpness of edges and spatial detail. In compressed video, it results from a loss of high frequency information during coding. Measurement of this artifact is based on the cosine of the angle between perpendiculars to planes in adjacent pixels which is a good characteristic of picture smoothness [17].”, Section IV Subsection A Paragraph 3 – “Ringing artifacts are visible for all compression techniques, especially when the image or video is transformed into a frequency domain. Ringing is a spurious reconstruction of pixel values. It is more evident along high contrast edges, especially if the edges are in areas with a generally smooth texture [18]. The noisiness metric estimates the noise level by a local variance of flat areas.” – teaches wherein the characteristic information includes, in a case where the input data is the image data, a value related to a blur amount of the image data (calculates blurring metrics based on cosine of angle between perpendiculars to planes in adjacent pixels) and noise according to sensitivity at the time of capturing the image data (determines noisiness metric that estimates the noise level by a local variance of flat areas, thus including characteristic information related to an amount of noise according to sensitivity at the time of capturing image data)), and
in a case where the input data is the sound data, first characteristic information related to noise sound and second characteristic information related to a sound volume (Leszczuk, Section IV Subsection C Paragraph 1 – “Clipping– the original audio signal can be clipped in certain situations during the recording due to environmental noise or recording equipment. The metric for this artifact is based on the detection of consecutive high levels of signal.” and in Section IV Subsection C Paragraph 2 – “Mute– signal losses are one of the most common degradation in audio streaming at low bit rates. The detection algorithm is based on two different thresholds, once the signal has frequency components only in the hearing range (between 20 Hz and 20 kHz). These thresholds are: the minimum level of signal noticeable by the human ear, and the duration of the shortest silent interval perceptible as a differentiated mute. The value of these thresholds must be adaptive to the nature of the sound, as experimental test have demonstrated.” – teaches the evaluation information and the characteristic information include, in a case where the input data is the sound data, a value related to a sound volume (detects muting based on two different thresholds, where the thresholds are the minimum level of signal and duration of shortest silent interval, thus determining a value related to sound volume) and noise sound (calculates clipping metric based on detection of consecutive high levels of signal, thus including characteristic information related to noise sound))
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the characteristic information of Leszczuk to the input data, substitute data, and decision unit of Zhao. Doing so would take the nature and characteristics of the video and audio artifacts into consideration for quality evaluations (Leszczuk Section IV Paragraph 1).
Regarding claim 12, the combination of Zhao and Leszczuk teaches the apparatus according to claim 5, wherein
the plurality of pieces of input data includes a plurality of pieces of image data that have been captured consecutively, the evaluation information is a value corresponding to a luminance difference between a piece of image data succeeding in a time series and a piece of image data preceding in the time series (Leszczuk, Section IV Subsection A Paragraph 10 – “Contrast is also stated as an important artifact for image/video quality assessment [22]–[24]. Contrast is the difference in luminance and/or color that makes an object (or its representation in an image or display) distinguishable.” – teaches the plurality of pieces of input data includes a plurality of pieces of image data that have been captured consecutively (determines contrast for video quality assessment, thus the plurality of pieces of input data include a plurality of image data captured consecutively in a video), the evaluation information is a value corresponding to a luminance difference between a piece of image data succeeding in a time series and a piece of image data preceding in the time series (determines a luminance difference corresponding to a video for video quality assessment. Thus teaching determining a luminance difference between a piece of image data succeeding in a video and a piece of image data preceding in the video)), and
the characteristic information is a value related to a noise amount at the time of capturing a respective piece of image data (Leszczuk, Section IV Subsection C Paragraph 1 – “Clipping– the original audio signal can be clipped in certain situations during the recording due to environmental noise or recording equipment. The metric for this artifact is based on the detection of consecutive high levels of signal.” – teaches wherein the evaluation information is a value related to a noise amount at the time of capturing the respective piece of image data (determines metric of clipping artifact based on detection of consecutive high levels of signal during recording, thus determining a noise amount at the time of capturing, or recording, a respective piece of image data)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the characteristic information of Leszczuk to the data and apparatus of Zhao. Doing so would take the nature and characteristics of the video and audio artifacts into consideration for quality evaluations (Leszczuk Section IV Paragraph 1).
Regarding claim 13, the combination of Zhao and Leszczuk teaches the apparatus according to claim 5, wherein
the plurality of pieces of input data include three or more types of data including image data and sound data (Zhao, Section 3 Paragraph 1 – “Given a set of video segments S, we use x= (xa, xv, xt) to represent the raw multimodal features for a video segment s∈S, where xa, xv and xt represent the raw features of acoustic, visual and textual modalities respectively.” – teaches wherein the plurality of pieces of input data include three or more types of data including image and sound data (data includes three or more types of data including visual, acoustic, and textual modalities)).
Regarding claim 14, the combination of Zhao and Leszczuk teaches the apparatus according to claim 13, wherein
the decision unit determines, for each piece of plurality of input data, a suitability of that input data by comparing evaluation information calculated for that input data with a first threshold (Leszczuk, Fig. 1 and Section IV Subsections A & C – teaches determining, for each of a plurality of input data, a suitability of that input data by comparing evaluation information calculated for that input data with a first threshold (determines metrics of video and audio artifacts such as blurring, exposure time distortion, clipping, and muting and compares the evaluation information with a threshold, and thus determining a suitability of that input data by comparing evaluation information for that input data with a first threshold)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the calculations of evaluation information and predetermined thresholds of Leszczuk to the inference model and decision unit of Zhao. Doing so would provide automatic quality measurements associated with occurrences of perceptible degradation in audio/video (Leszczuk, Introduction)
Claim(s) 4 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao and Leszczuk as applied to claims 1 and 19-20 above, and further in view of Nasrollahi et al. (NPL: Face Quality Assessment in Video Sequences, published 2008, hereinafter “Nasrollahi”).
Regarding claim 4, the combination of Zhao and Leszczuk teaches the apparatus according to claim 2.
The combination of Zhao and Leszczuk fails to explicitly teach wherein in a case where all of the plurality of pieces of input data are determined to be unsuitable by the decision unit, the decision unit changes a result of determination of a piece of input data whose evaluation information is the highest to suitable.
However, analogous to the field of the claimed invention, Nasrollahi teaches: wherein
in a case where all of the plurality of pieces of input data are determined to be unsuitable by the decision unit, the decision unit changes a result of determination of a piece of input data whose evaluation information is the highest to suitable (Nasrollahi, Section 3.5 Paragraph 1 – “After calculating the four above mentioned features for each of the images in a given sequence, we combine the scores of these features into a general score for each image, as shown in the following equation: Eq. (9) … The images are sorted based on their combined scores and depending on the application, one or more images with the greatest values in S are considered as the highest quality image(s) in the given sequence.” and in Section 4 Paragraph 4 – “By the way, even for the poor quality images, although it is possible that the images be sorted in different way by the system and the human, but in 100% of the cases we can find the best chosen image by the human inside the first four chosen images by the system.” – teaches wherein in case where all the plurality of pieces of input data are determined to be unsuitable (in a case where all images are poor quality), the decision unit changes a result of determination of a piece of input data whose evaluation information is the highest to suitable (even if all images are poor quality, selects best image in 100% of the cases, where best image selection is based on images with highest general score)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the evaluation information and determination of suitable data in a case where all input data are unsuitable of Nasrollahi to the decision unit, suitability determination, and evaluation information of Zhao and Leszczuk. Doing so would provide mechanisms for choosing the best image in terms of quality in a sequence of images (Nasrollahi, Introduction).
Regarding claim 11, the combination of Zhao and Leszczuk teaches the apparatus according to claim 5, wherein
the characteristic information is a value related to a noise amount at the time of capturing a respective piece of image data (Leszczuk, Section IV Subsection C Paragraph 1 – “Clipping– the original audio signal can be clipped in certain situations during the recording due to environmental noise or recording equipment. The metric for this artifact is based on the detection of consecutive high levels of signal.” – teaches wherein the characteristic information is a value related to a noise amount at the time of capturing the respective piece of image data (determines metric of clipping artifact based on detection of consecutive high levels of signal during recording, thus determining a noise amount at the time of capturing, or recording, a respective piece of image data)).
The combination of Zhao and Leszczuk fails to explicitly teach the plurality of pieces of input data include a plurality of pieces of image data in which an orientation of a subject is different and the evaluation information is a value corresponding to an orientation of a subject of a respective piece of image data.
However, analogous to the field of the claimed invention, Nasrollahi teaches:
the plurality of pieces of input data include a plurality of pieces of image data in which an orientation of a subject is different (Nasrollahi, Fig. 4 – teaches wherein the plurality of pieces of input data include a plurality of pieces of image data in which an orientation of a subject is different (Fig. 4 shows a sequence of different head poses and associated values)),
the evaluation information is a value corresponding to an orientation of a subject of a respective piece of image data (Nasrollahi, Fig. 4 and in Section 3.1 Paragraphs 3-4 –“Now we calculate the distance between these two centers as: Eq. (3)”, “The minimum value of this distance in a sequence of images gives us the least out-of-plan rotated face as shown in figure 4. To convert this value to a local score in that sequence we use the following equation for each of the images in the sequence: Eq. (4)” – teaches the evaluation information is a value corresponding to an orientation of a subject of a respective piece of image data (determines a distance and a score corresponding to an orientation of a subject, or a rotated face, of a respective piece of image data))
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the image data in which an orientation of a subject is different and evaluation information corresponding to orientation of a subject of Nasrollahi to the data, evaluation information, and apparatus of Zhao and Leszczuk. Doing so would provide mechanisms for choosing the best image in terms of quality in a sequence of images (Nasrollahi, Introduction) and mechanisms for scoring images based on head pose (Nasrollahi, Section 3.1).
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao and Leszczuk as applied to claims 1 and 19-20 above, and further in view of Nookula et al. (US Patent No. 11,544,577, published Jan. 2023, hereinafter “Nookula”).
Regarding claim 15, the combination of Zhao and Leszczuk teaches the apparatus according to claim 14, wherein
the decision unit determines a suitability of input data according to a result of determination based on the first threshold (Leszczuk, Fig. 1 and Section IV Subsections A & C – teaches determining a suitability of input data according to a result of determination based on the first threshold (determines metrics of video and audio artifacts such as blurring, exposure time distortion, clipping, and muting and compares the evaluation information with a threshold, and thus determining a suitability of that input data by comparing evaluation information for that input data with a first threshold))
The combination of Zhao and Leszczuk fails to explicitly teach the decision unit obtains a difference between a highest piece of evaluation information and a next highest piece of evaluation information and, when it is determined that the difference is greater than or equal to a second threshold, determines a suitability of input data and, when it is determined that the difference is less than the second threshold, determines that all pieces of input data are suitable regardless of the result of determination based on the first threshold.
However, analogous to the field of the claimed invention, Nookula teaches:
the decision unit obtains a difference between a highest piece of evaluation information and a next highest piece of evaluation information and, when it is determined that the difference is greater than or equal to a second threshold, determines a suitability of input data and, when it is determined that the difference is less than the second threshold, determines that all pieces of input data are suitable regardless of the result of determination based on the first threshold (Nookula, Pg. 11 Col. 2 Lines 21-29 – “In some embodiments, the filter can be a differential-type filter that generates difference representations between consecutive elements of a data stream to determine which elements are to be passed on to be used as inputs for the ML model. In some embodiments, the filter can be a “smart” filter utilizing a ML-type model such as a neural network that can be trained using outputs from the ML model, allowing the filter to “learn” which elements of the data stream are the most likely to be of value to be passed on.” Pg. 13 Col. 5 Lines 26-41 – “For example, in an embodiment where the input data stream 110 is a video stream, every element could be an image that can be represented in a matrix of vectors. Thus, difference generation unit 202 (e.g., implemented as a software module, hardware module, or combination thereof) of the differential filter 106A can perform a difference operation between a first frame 200A and a next frame 200B to find a difference 204 between the two frames. An analysis unit 206 (e.g., implemented as a software module, hardware module, or combination thereof) of the differential filter 106A can determine, at decision block 208, whether the difference 204 is higher than (e.g., meets or exceeds) a sensitivity threshold value 220 (that could be defined by a user 222). If so, the analysis unit 206 can determine that the electronic device 102 is to send the later frame on to be used as an input for the ML model 108 (at block 212) for inference; otherwise, the analysis unit 206 may halt processing of the particular pair of data stream elements (e.g., frames 200A-200B) and continue by analyzing another pair of consecutive frames” – teaches obtaining a difference (determines difference of a representation between elements) between a highest piece of evaluation information and a next highest piece of evaluation information (learns elements most likely to be of value, thus determining elements with highest evaluation information), and when it is determined that the difference is greater than or equal to a second threshold, determining a suitability of input data and, when it is determined that the difference is less than a second threshold, determines that all pieces of input data are suitable regardless of the result of determination based on the first threshold (determines difference of a representation between elements of data stream to determine which elements are to be passed on as inputs, and if difference meets threshold, determines that all elements are suitable for an ML model. If difference does not meet threshold, then determines elements are not suitable for ML model, thus determining a suitability of input data if the difference does not meet the threshold)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the difference determination and second threshold of Nookula to the first threshold, suitability determinations, and evaluation information of Zhao and Leszczuk. Doing so would provide mechanisms that generate representations between elements of a data stream to determine which elements are to be passed on to be used as inputs for the ML model (Nookula, Pg. 11 Col. 2).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Antonides et al. (US Pub. No. 2022/0405578, published Dec. 2022) teaches systems and methods for multi-modal fusion. Teaches learning one or more latent spaces using at least two different types of data modalities. Teaches replacing data of a certain modality using generated data, depending on a threshold value.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOUIS C NYE whose telephone number is 571-272-0636. The examiner can normally be reached Monday - Friday 9:00AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOUIS CHRISTOPHER NYE/Examiner, Art Unit 2141
/TAN H TRAN/Primary Examiner, Art Unit 2141