DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on 04/21/2026 . Claims 1-3, 5-12 and 14 are pending and have been examined.
Any previous objection/rejection not mentioned in this Office Action has been withdrawn by the Examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Change in Examiner
The Examiner of Record has changed from Cameron Young to SPE Paras Shah.
Response to Amendments/Arguments
With respect to the 35 USC 103 rejections, the Applicant asserts that the claims have been amended to explicitly recite that the watermark of the claims is based on the feature vector and the private key but that the feature vector and the private key are not equivalent to the watermark. However, a new rejection under 35 USC 112(a) has been issued as the Examiner is not able to locate support as to how the watermark is not equivalent to the feature vector and the private key. (please see below rejection for full details). The Examiner also notes that the same amendments have not been made to claim 12. Therefore, the prior art rejections of claim 12 have been maintained and the prior art rejections have been removed from claim 1 and its dependents.
With respect to the arguments on page 7, where Chalamala does not specifically disclose or suggest “generating a private key based on an encryption of the feature vector using an encryption algorithm”. The Applicant then refers to the amendments which were made to the claims. However, as noted above, these amendments were only made to claim 1 and not claim 12. Therefore, the rejections for claim 12 are still maintained as no separate arguments are presented for this set of claims.
With respect to the in re keller arguments, the Examiner appreciates the explanations provided. However, there does not appear to be a specific argument to respond to and therefore has only been acknowledged by the Examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
“a voice collection unit” in claim 1,
“a frame generation unit” in claim 1,
“a frequency analysis unit” in claims 1 and 5,
“a neural network learning unit” in claim 1,
“a watermark generation unit” in claim 6,
“a watermark embedment unit” in claims 6 - 8,
“a watermark extraction unit” in claim 6,
“an encryption generation unit” in claim 9,
“an authentication comparison unit” in claim 9,
and “an authentication determination unit” in claim 9.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-3 and 5-11 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. The Applicant has added the limitation of “wherein the feature vector and the private key are not equivalent to the watermark”. However, the Examiner asserts that there is no support found that shows the watermark to be non equivalent to the feature vector and private key. Several passages of the Applicant’s as filed specification support this conclusion. For example, paragraph [0066]-[0067], [0081] clearly describes that the watermark is generated based on the private key, where the private key is generated from the encryption of the feature vector. However, nowhere in these paragraphs is it made clear that the feature vector and privacy key is non equivalent to the watermark. In fact, these paragraphs appear to suggest the watermark can be equivalent to the “primary key’ where the primary key is determined from encrypted feature vectors. If the Applicant wishes to challenge this finding, the Applicant must show support from the Specification/Drawings of how the limitation is supported as it is presented.
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Chalamala in view of Huffman in further view of U.S. Patent Application Publication No. 2015/0325246 A1 to Chi-Man Pun et al. (hereinafter Pun) and in further view of U.S. Patent Application Publication No. 2020/0035247 A1 to Constantine T. Boyadjiev et al. (hereinafter Boyadjiev).
Regarding claim 12, Chalamala teaches a voice authentication method comprising: (Chalamala teaches a multi-factor authentication system that performs a method of authentication including a voice (i.e., voice authentication method). Chalamala at ¶ [0010].)
a voice collection step of collecting voice information obtained by digitizing a speaker's voice; (Chalamala teaches a user interface comprising a microphone in communication with a computer processor that collects audio for use in voice authentication. (i.e., the audio received by the microphone is processed by the compute processor, thus the audio is digitized in order to be processed or is digitally captured by the microphone.) Chalamala at ¶ [0030].)
Chalamala, however, does not teach a learning model step of generating a voice image based on the collected voice information of the speaker, causing a deep neural network (DNN) model to learn the voice image, and extracting a feature vector for the voice image.
In a similar field of endeavor (e.g., generating and watermarking audio/speech.) Huffman teaches a learning model step of generating a voice image based on the collected voice information of the speaker, causing a deep neural network (DNN) model to learn the voice image, and extracting a feature vector for the voice image. (Huffman teaches extracting features from voice data in the vector space (i.e., extracting a feature vector) using a machine learning system. (Huffman at ¶ [0045].) Further, Huffman teaches a watermark machine learning system may also be referred to as a watermark network which may be a deep neural network. (i.e., the machine learning system may be a deep neural network for "learning" the voice image and extracting the feature vector.) Huffman at ¶¶ [0048] - [0049]. Further, Huffman teaches the voice data being a spectrogram (i.e., a voice image). Huffman at ¶ [0063] and Fig. 5. Thus, Huffman teaches representing the speech data as a spectrogram (i.e., generating a voice image based on the speech data. In order to represent speech data as a spectrogram, the spectrogram must be generated/created), a deep neural network learning the voice image (a machine learning system (i.e., a deep neural network) processing the speech data), and extracting features from the speech data in the vector space using the machine learning system. (i.e., extracting a feature vector from the speech data.))
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Chalamala with the teachings of Huffman to provide a learning model server configured to generate a voice image based on the collected voice information of the speaker, cause a deep neural network (DNN) model to learn the voice image, and extract a feature vector for the voice image. Doing so would have improved optimization of watermarked audio by hiding the watermark better than traditional methods as recognized by Huffman at ¶ [0126].
Further, the combination of Chalamala-Huffman teaches an encryption generation step of encrypting the feature vector to generate a private key corresponding to the feature vector using an encryption algorithm; (Chalamala teaches generating an encrypted passphrase and embedding the encrypted passphrase into an audio signal. (i.e., generating a private key.) Chalamala at ¶ [0033] – [0034]. Further, a person of ordinary skill in the art would have understood that encryption of a feature vector requires the use of some form of encryption algorithm. Chalamala teaches encryption of a passphrase, as laid out above, and as such Chalamala teaches encrypting using an encryption algorithm.)
a watermark generation step of generating and storing a watermark and individual information based on the private key and the feature vector; (Chalamala teaches generating a plurality of encryption parameters and embedding them in the audio signal as a watermark to generate a watermarked audio signal. (i.e., generating and storing the watermark.) Chalamala at ¶¶ [0033] – [0034]. Further, Chalamala’s generation and embedding of a authentication parameters (i.e., a private key) in view of Huffman’s extraction of a feature vector for embedding into an image as a watermark would have been obvious to combine to encrypt the private key based on the extracted feature vector as both elements are the target of the embedding, therefore the limitations of the claims can be achieved simply adding Huffman’s extraction of a feature vector to Chalamala’s generation of authentication parameters and embedding thereof. Chalamala at ¶¶ [0033] – [0034] and Huffman ¶¶ [0045] – [0049].)
a watermark embedment step of embedding the watermark and the individual information …; (Chalamala teaches embedding the encryption parameters in an audio signal as a watermark. (i.e., embedding the watermark in voice image or voice conversion data.) Chalamala at ¶¶ [0033] – [0034].)
an authentication comparison step of comparing the sameness between the encrypted feature vector and a feature vector of an authentication target using the private key; (Chalamala teaches comparing encrypted data with other encrypted data to authenticate a speaker. (i.e., perform a sameness comparison of encrypted data (feature vector) and other encrypted data (another feature vector)) Chalamala at ¶ [0050]. Further, the encrypted data in Chalamala is an encrypted passphrase (i.e., private key) therefore the comparison is using the private key. Chalamala ¶ [0050] and ¶¶ [0033] – [0034].)
an authentication determination step of determining whether authentication is successful for the speaker based on a comparison result, and determining whether to extract the watermark and the individual information; (Chalamala teaches determining if authentication was successful and subsequently removing the watermark from the audio signal as a result of successful authentication. Chalamala at ¶ [0026])
and a watermark extraction step of extracting the watermark and the individual information that have been pre-stored based on an authentication result. (Chalamala teaches removing the watermark from the voice-image as a result of speaker authentication. Chalamala at ¶ [0026].)
Chalamala in view of Huffman, however, does not teach embedding the watermark and the individual information into a pixel of the voice image or voice conversion data.
Pun teaches embedding the watermark and the individual information into a pixel of the voice image or voice conversion data. (Pun teaches embedding information in specific pixels (i.e., the first n-1 pixels of the block.) of audio data. Pun at ¶¶ [0029] – [0036].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Chalamala-Huffman with the teachings of Pun to provide embedding the generated watermark and the individual information into a pixel of the voice image or the voice conversion data. Doing so would have either improved acoustic quality or achieved higher embedding space as recognized by Pun at ¶ [0055].
Chalamala-Huffman-Pun, however, do not teach “wherein the learning model step further includes: a frame generation step of generating a voice frame for a predetermined time based on the voice information; a frequency analysis step of analyzing a voice frequency based on the voice frame, and generating the voice image in time series by imaging the voice frequency; a neural network learning step of causing the deep neural network model to learn the voice image; and a feature vector extraction step of extracting the feature vector of the learned voice image.”
In a similar field of endeavor (e.g., machine learning processes for voice authentication.) Boyadjiev teaches wherein the learning model step further includes:
a frame generation step of generating a voice frame for a predetermined time based on the voice information; (Boyadjiev teaches generating a plurality of multi-dimensional feature vectors (i.e., voice images/ frames) based on input speech signals. Boyadjiev at ¶¶ [0061] – [0066] and Fig. 7. Boyadjiev teaches looking for a specific frequency over time that matches a control voice sample and searching for a specific portion of a voice sample and isolating the specific portion matching the voice control sample. Boyadjiev at ¶¶ [0061] - [0065]. As such, the specific portion matching the voice control sample would correspond to a specific time of the voice control sample and because the voice control sample is generating a voice frame for a predetermined time based on the voice information.)
a frequency analysis step of analyzing a voice frequency based on the voice frame, and generating the voice image in time series by imaging the voice frequency; (Boyadjiev teaches extracting the multi-dimensional feature vector (voice frame) based on frequency (i.e., a frequency analysis is performed to determine what to remove.). Boyadjiev at ¶ [0088]. Further, Boyadjiev teaches extracting multi-dimensional acoustic feature vectors (e.g., RGB multi-dimensional acoustic feature vector i.e., a voice image) and processing the acoustic feature vectors as frequency over time in the process of matching (i.e., the matching of voice features is done in time series) Boyadjiev at ¶¶ [0060] – [0066].)
a neural network learning step of causing the deep neural network model to learn the voice image; (Boyadjiev teaches extracting a feature vector using a neural network. (i.e., the neural networks extract the feature vector to obtain the multi-dimensional feature vector (i.e., extract information from the multi-dimensional feature vector.).) Boyadjiev at ¶ [0017].)
and a feature vector extraction step of extracting the feature vector of the learned voice image. (Boyadjiev teaches the extraction of acoustic feature vectors from the multi-dimensional feature vector using neural networks. (i.e., the neural networks are configured to extract the feature vector to “learn” the voice image (multi-dimensional feature vector).). Boyadjiev at ¶ [0017].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Chalamala-Huffman-Pun with the teachings of Boyadjiev (hereinafter Chalamala-Huffman-Pun-Boyadjiev) to provide the limitations of claim 12. Doing so would have improved accuracy of the feature vectors as recognized by Boyadjiev at ¶ [0060].
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Chalamala-Huffman-Pun-Boyadjiev as applied to claim 12 above, and further in view of Ross.
Regarding claim 14, Chalamala-Huffman-Pun-Boyadjiev teaches all the limitations of claim 12 as laid out above. Chalamala-Huffman-Pun-Boyadjiev, however, does not teach an authorization step of, when authentication is successful, granting access and modification authority to the extracted voice information and individual information; and a forgery warning step of, when authentication fails, outputting a warning signal for information forgery.
Ross teaches an authorization step of, when authentication is successful, granting access and modification authority to the extracted voice information and individual information; (Ross teaches granting access to a device to a user after successful authentication. (i.e., the user may now access the device. In the case of a computer accessing a device could be administrator privileges (i.e., editing secure information).). Ross at ¶ [0036].)
and a forgery warning step of, when authentication fails, outputting a warning signal for information forgery. (Ross teaches alerting (i.e., issuing a warning signal) to the user when an unauthorized user attempts to access a device. (i.e., if authentication fails, (unauthorized access) the user is alerted of the malicious attempt.) Ross at ¶ [0089].)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to combine the teachings of Chalamala-Huffman-Pun-Boyadjiev with the teachings of Ross to provide the limitations of claim 14. Doing so would have allowed for reliable implementation of authorization as recognized by Ross at ¶ [0042].
Allowable Subject Matter
Claims 1-3 and 5-11 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), 1st paragraph, set forth in this Office action.
No reasons for allowance are being provided due to the pending 35 USC 112(a) rejections.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Takeyasu et al. (JP 2015011580A) is cited to disclose creating a watermark to be included into a voice signal.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PARAS D SHAH whose telephone number is (571)270-1650. The examiner can normally be reached Monday-Thursday 7:30AM-2:30PM, 5PM-7PM (EST), Friday 8AM-noon (EST).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
07/31/2026