Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAIL ACTION
Priority
Acknowledgment is made of applicant's claim for foreign priority under 35 U.S.C. 119(a)-(d). The certified copy has been placed of record in the file.
Information Disclosure Statement
The information disclosure statement (IDS) was submitted on 3/4/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
CLAIM INTERPRETATION
4. The following is a quotation of 35 U.S.C. 112(f): (FP 7.30.03)
(f) ELEMENT IN CLAIM FOR A COMBINATION.—An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
5. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as "configured to" or "so that"; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
6. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: an acquisition unit which acquires a frame; a streaming feature generation unit which generates a first feature; a streaming character generation unit which generates a first character; a non-streaming feature generation unit which generates a second feature sequence; a streaming character generation unit which generates a second character string; and a learning unit which performs Knowledge Distillation in claim 1.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. (FP 7.30.06)
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-6 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claimed invention is directed to non-statutory subject matter because the claim(s) as a whole, considering all claim elements both individually and in combination, do not amount to significantly more than an abstract idea. As summarized in the 2019 Revised Patent Subject Matter Eligibility Guidance, examiners must perform a Two-Part Analysis for Judicial Exceptions.
Step 1
In Step 1, it must be determined whether the claimed invention is directed to a process, machine, manufacture or composition of matter. The instant invention encompasses three sets of claims: a device in claims 1-4 (i.e., a manufacture), a method in claim 5 (i.e., a process) and a non-transient storage medium in claim 6 (i.e., a manufacture). All claims are directed to one of the four statutory categories and meet the requirements of step 1.
Step 2A
Prong One
The claimed invention is directed to an abstract idea without significant more. The instant invention is broadly directed to “generating character from stream/non-stream encoder and decoder”. Claim 1 recites the following (with emphasis added):
Claim 1: A voice recognition device, comprising:
an acquisition unit which acquires a frame per unit time of a voice stream;
a streaming feature generation unit which generates a first feature from the frame using a streaming encoder;
a streaming character generation unit which generates a first character from the first feature using a streaming decoder;
a non-streaming feature generation unit which generates a second feature sequence from a first feature sequence obtained by joining the first feature of each of the plurality of frames using a non-streaming encoder;
a streaming character generation unit which generates a second character string from the second feature sequence using a plurality of non-streaming decoders; and
a learning unit which performs Knowledge Distillation between the streaming encoder and the non-streaming encoder on the basis of the first feature sequence and the second feature sequence.
The bold portions of claim 1 encompass the abstract idea, which is also encompassed by the dependent claims 2-4, and substantially also encompassed by claims 5 and 6.
Claims 1, 5 and 6 recite the steps to generate feature and character from input voice stream then performing Knowledge Distillation based on the generated features. These limitations, when given their broadest reasonable interpretation, are directed to certain performing of organizing human activity and mental processes, which is abstract idea.
Prong Two
This judicial exception is not integrated into a practical application because mere instruction to implement on computers (i.e. storage medium or computer in claim 6) or a computer model (units here in claim 1), or merely using computers or processor as a tool to perform the abstract idea, adding insignificant extra solution activity, and/or generally linking the use of the abstract idea to a technological environment for field of use is not considered integration into a practical application. Claim 1 recites using natural language prompt to generate output data of the trained natural language model. Using input voice stream to generate features and perform Knowledge Distillation based on the generated features is a generic feature of audio signal process, which does not represent a technological improvement. The using of the computer/processor and general audio signal processing units does not add improvement to the functioning of a computer or to any other technology field, which failed to enable the abstract idea to integrate into a practical application. The claims are drafted in a result-oriented fashion, without the requisite specificity needed to provide a nonabstract technological solution. The computing system and voice signal processing units are directed to the components of a system amount to merely field of use type limitations and/or extra solution activity to implement the abstract idea as presented.
Step 2B
Step 2B in the analysis requires us to determine whether the claims do significantly more than simply describe that abstract method. Mayo, 132 S. Ct. at 1297. We must examine the limitations of the claims to determine whether the claims contain an "inventive concept" to "transform" the claimed abstract idea into patent-eligible subject matter. Alice, 134 S. Ct. at 2357 (quoting Mayo, 132 S. Ct. at 1294, 1298). The transformation of an abstract idea into patent-eligible subject matter "requires 'more than simply stat[ing] the [abstract idea] while adding the words 'apply it."' Id. (quoting Mayo, 132 S. Ct. at 1294) (alterations in original). "A claim that recites an abstract idea must include 'additional features' to ensure 'that the [claim] is more than a drafting effort designed to monopolize the [abstract idea].'" Id. (quoting Mayo, 132 S. Ct. at 1297) (alterations in original). Those "additional features" must be more than "well-understood, routine, conventional activity." Mayo, 132 S. Ct. at 1298.
The present claims include the additional elements other than the abstract idea which include a processor, storage medium, voice processing units (in claim 1 and 6). These additional elements are merely conventional computer and computer model. Any potentially technical aspects of the claims are well-known generic computer components performing conventional functions (e.g., a processor performing a mental process). The present claims have been analyzed both individually and in combination and, the instant claims do not provide any improvement of the functioning of the computer or improvement to computer technology or any other technical field. There do not appear to be any meaningful limitations other than those that are well-understood, routine and conventional in the field. Thus, the present claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claims 1-4 are not patent eligible.
Claims 5 and 6 recite similar limitations of claims 1-4, thus are abstract idea and not patent eligible.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
35 U.S.C. 101 requires that a claimed invention must fall within one of the four eligible categories of invention (i.e. process, machine, manufacture, or composition of matter) and must not be directed to subject matter encompassing a judicially recognized exception as interpreted by the courts. MPEP 2106. The four eligible categories of invention include: (1) process which is an act, or a series of acts or steps, (2) machine which is an concrete thing, consisting of parts, or of certain devices and combination of devices, (3) manufacture which is an article produced from raw or prepared materials by giving to these materials new forms, qualities, properties, or combinations, whether by hand labor or by machinery, and (4) composition of matter which is all compositions of two or more substances and all composite articles, whether they be the results of chemical union, or of mechanical mixture, or whether they be gases, fluids, powders or solids. MPEP 2106(I)
Claims 1-4 are rejected under 35 U.S.C. 101 as not falling within one of the four statutory categories of invention because the claimed invention is directed to computer program per se. See MPEP 2106(I). A claim directed toward a non-transitory computer-readable medium having the program encoded thereon establishes a sufficient functional relationship between the program and a computer so as to remove it from the realm of “program per se”. MPEP 2111.05(III). Hence, adding the limitation of “any of the units or encoder/decoder are implemented by a hardware (e.g. a processor , a computer, or the software units (modules) are stored in a non-transitory computer-readable medium” would resolve this issue.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al (US 20230335126 A1) in view of CHEN (CN 117456986 A).
Regarding claim 1, Huang discloses a voice recognition device [e.g. FIG. 1-2; automatic speech recognition (ASR) systems], comprising:
an acquisition unit which acquires a frame per unit time of a voice stream [e.g. FIG. 1; microphone 16 capturing voice stream, a sequence of acoustic frames; [0031-0034]; stream during time 1];
a streaming feature generation unit which generates a first feature from the frame using a streaming encoder [e.g. FIG. 2; causal encoder to generate a first higher order feature representation 212 for a corresponding acoustic frame 110 in the sequence of acoustic frames];
a streaming character generation unit which generates a first character from the first feature using a streaming decoder [e.g. decoder 206 to generate the output Yr of the decoder 206 represents the output of the Softmax layer];
a non-streaming feature generation unit which generates a second feature sequence from a first feature sequence obtained by joining the first feature of each of the plurality of frames using a non-streaming encoder [e.g. the second encoder 220 generates the second higher order feature representations 222 using only the first higher order feature representation 212 as input];
a streaming character generation unit which generates a second character string from the second feature sequence using a plurality of non-streaming decoders [e.g. FIG. 2; [0049]; In the non-streaming mode, the decoder 206 uses the joint network 230 to combine the first higher order feature representation 212 and the second higher order feature representation 222 output by the cascading encoder 204, as well as the average embedding 242 generated by the prediction network 240 to generate an initial transcription (e.g., decoder output) 232]; and
Although Huang disclose a learning unit to perform speech recognition [e.g. training LM 160 and 170M] on the basis of the first feature sequence and the second feature sequence [e.g. FIG. 2]; it is noted that Huang differs to the present invention in that Huang fails to explicitly disclose the details of the learning unit.
However, CHEN teaches the well-known concept of a learning unit [e.g. FIG. 1-3; training techniques for ASR models] which performs Knowledge Distillation [e.g. FIG. 2-3; page 3-4; content of the invention section; determining knowledge distillation loss] between the streaming encoder [e.g. FIG. 4; a stream encoder for streaming speech recognition model] and the non-streaming encoder [e.g. a non-stream encoder for non-streaming speech recognition model] on the basis of the first feature sequence and the second feature sequence [FIG. 3-4; output of the first decoder for streaming speech recognition model and the output for the second decoder for non-streaming speech recognition model].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the speech recognition system disclosed by Huang to exploit the well-known determining Knowledge Distillation loss technique taught by CHEN as above, in order to provide the low-delay flow type recognition and guarantee the recognition quality [See CHEN; abstract].
Regarding claim 2, Huang and CHEN further disclose the learning unit performs the Knowledge Distillation between the streaming encoder and the non- streaming encoder so that the first feature sequence is similar to the second feature sequence [e.g. CHEN: FIG. 2-3; S203; page 9-10].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the speech recognition system disclosed by Huang to exploit the well-known determining Knowledge Distillation loss technique taught by CHEN as above, in order to provide the low-delay flow type recognition and guarantee the recognition quality [See CHEN; abstract].
Regarding claim 3, Huang and CHEN further disclose the learning unit further performs the Knowledge Distillation between the streaming decoder and the plurality of non-streaming decoders on the basis of a likelihood of a first character string in which the first characters generated from each of the plurality of first features are arranged in chronological order and a likelihood of the second character string generated from the second feature sequence [e.g. CHEN: FIG. 2-3; S203; page 9-10; S204, when reaching the preset training condition, stopping training, obtaining the trained stream voice identification model; the iteration times exceeds the limit value, the distillation training process is ended].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the speech recognition system disclosed by Huang to exploit the well-known determining Knowledge Distillation loss technique taught by CHEN as above, in order to provide the low-delay flow type recognition and guarantee the recognition quality [See CHEN; abstract].
Regarding claim 4, Huang and CHEN further disclose the streaming decoder includes at least a predictor configured to predict the first character [e.g. Huang: FIG. 2; predictor], each of the plurality of non-streaming decoders includes at least an attention decoder that is a decoder including an attention mechanism [e.g. CHEN: FIG. 2-3; attention decoder], and the learning unit performs the Knowledge Distillation between the predictor and the attention decoder so that a likelihood of the first character string in which the first character predicted using the predictor is arranged in chronological order is similar to a likelihood of the second character string output using the attention decoder [e.g. CHEN: FIG. 2-3; S203; page 9-10; S204, when reaching the preset training condition, stopping training, obtaining the trained stream voice identification model; the iteration times exceeds the limit value, the distillation training process is ended].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the speech recognition system disclosed by Huang to exploit the well-known determining Knowledge Distillation loss technique taught by CHEN as above, in order to provide the low-delay flow type recognition and guarantee the recognition quality [See CHEN; abstract].
Regarding claim 5, this is a method that includes same limitation as in claim 1 above, the rejection of which are incorporated herein.
Regarding claim 6, this is a non-transitory computer-readable storage medium that includes same limitation as in claim 1 above, the rejection of which are incorporated herein.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Tripathi et al (US 20240177706 A1).
Kim et al (US 20220319501 A1).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHUBING REN whose telephone number is (571)272-2788. The examiner can normally be reached Monday-Friday 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZHUBING REN/ Primary Examiner, Art Unit 2658