DETAILED ACTION
Claims 1 – 20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
EXAMINER’S COMMENT
Regarding claims 17 – 20 which are directed to ‘[o]ne or more hardware storage devices,’ the Examiner considered [00100] of the Specification where it provides that ‘[c]omputer readable media (e.g. hardware storage device(s) 140 of Fig. 1) that store computer-executable instructions (e.g., computer-readable instructions 118 of Fig. 1) are physical hardware storage media/devices that exclude transmission media.’ Based on this, the claimed hardware storage device do not include transitory storage devices, and are hereby not given a 35 U.S.C. 101 computer-readable medium rejection.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Instant claim 1 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent No. 12,277,927 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed to accessing training data that comprise audio spoken in a first language and text of the spoken audio being provided in a second language, applying an automatic speech translation model to receive and encode the audio data to generate translated language based on transcription label output, training the automatic speech translation model by applying the indicated training dataset, and generating a transcription in the second language of input audio data in the first language based on the trained end-to-end AST model.
Instant claim 9 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 7 of U.S. Patent No. 12,277,927 B2 (claim 7 depends on claim 6). Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed to accessing training data that comprise audio spoken in a first language and text of the spoken audio being provided in a second language, applying an automatic speech translation model to receive and encode the audio data to generate translated language based on transcription label output, training the automatic speech translation model by applying the indicated training dataset, and generating a transcription in the second language of input audio data in the first language based on the trained end-to-end AST model.
Instant claim 17 is rejected on the ground of nonstatutory double patenting as being obviously unpatentable over claim 1 of U.S. Patent No. 12,277,927 B2 (instant claim 17 is directed to one or more hardware storage devices while claim 1 of U.S. 12,277,927 B2 is directed to a method). Although the claims at issue are not identical, they are not patentably distinct from each other because they are both directed to accessing training data that comprise audio spoken in a first language and text of the spoken audio being provided in a second language, applying an automatic speech translation model to receive and encode the audio data to generate translated language based on transcription label output, training the automatic speech translation model by applying the indicated training dataset, and generating a transcription in the second language of input audio data in the first language based on the trained end-to-end AST model. One of ordinary skill in the art would have found it obvious to modify the techniques provided by the method of claim 1 (U.S. 12,277,927 B2), to arrive at the one or more hardware storage devices of instant claim 17, given the predictable result of making the techniques available to be loaded and executed on different computing devices without having to individually reprogramme the required instructions.
Instant claim
U.S. Patent 12,277,927 B2
Claim 1
Claim 1
A method comprising:
A method for implementing an end-to-end automatic speech translation (AST) model with a neural transducer, the method comprising:
accessing a training dataset comprising (i) an audio dataset comprising a spoken language utterance spoken in a first language and (ii) a text dataset comprising a transcription label that is written in a second language, the transcription label corresponding to the spoken language utterance;
accessing a training dataset comprising an audio dataset comprising spoken language utterances in a first language and a text dataset comprising transcription labels in a second language, the transcription labels corresponding to the spoken language utterances;
accessing an automatic speech translation (AST) model that is configured to receive and encode audio data, the AST model being further configured to generate translated language based on transcription label output;
accessing an end-to-end AST model based on a neural transducer comprising at least an acoustic encoder which is configured to receive and encode audio data, a prediction network which is integrated in a parallel model architecture with the acoustic encoder in the end-to-end AST model and configured to predict a subsequent language token based on a previous transcription label output;
training the AST model by applying the training dataset to the AST model;
applying the training dataset to the end-to-end AST model;
accessing input audio data that is in the first language; and
using the AST model to generate a transcription of the input audio data, the transcription of the input audio data being in the second language.
generating a transcription in the second language of input audio data in the first language based on the trained end-to-end AST model.
Claim 9
Claim 7
A computer system comprising:
one or more processors; and
A computer system comprising:
one or more processors; and (claim 6)
one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computer system to:
one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computer system to execute an end-to-end automatic speech translation (AST) model based on a neural transducer configured to receive input audio data in a first language and to generate a transcription of the input audio data in a second language, wherein the end-to-end AST model comprises: (claim 6)
access a training dataset comprising (i) an audio dataset comprising a spoken language utterance spoken in a first language and (ii) a text dataset comprising a transcription label that is written in a second language, the transcription label corresponding to the spoken language utterance;
The computer system of claim 6, wherein the end-to-end AST model is trained on a training dataset that comprises: an audio dataset comprising spoken language utterances in a first language and a text dataset comprising transcription labels in a second language, the transcription labels corresponding to the spoken language utterances. (Claim 7)
access an automatic speech translation (AST) model that is configured to receive and encode audio data, the AST model being further configured to generate translated language based on transcription label output;
an acoustic encoder comprising a plurality of temporal processing paths configured to receive and encode input audio data that comprises a particular number of frames which is configured to be separated into different sets of frames, wherein each temporal processing path is configured to process the particular number of frames according to a particular combination of one or more different sets of frames included in the input audio data, wherein the acoustic encoder is configured to output an intermediary transcription label for each different set of frames; (claim 6)
train the AST model by applying the training dataset to the AST model;
The computer system of claim 6, wherein the end-to-end AST model is trained on a training dataset (claim 7)
access input audio data that is in the first language; and
use the AST model to generate a transcription of the input audio data, the transcription of the input audio data being in the second language.
one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computer system to execute an end-to-end automatic speech translation (AST) model based on a neural transducer configured to receive input audio data in a first language and to generate a transcription of the input audio data in a second language, wherein the end-to-end AST model (claim 6).
Claim 17
Claim 1
One or more hardware storage devices that store instructions that are executable by one or more processors to cause the one or more processors to:
access a training dataset comprising (i) an audio dataset comprising a spoken language utterance spoken in a first language and (ii) a text dataset comprising a transcription label that is written in a second language, the transcription label corresponding to the spoken language utterance;
A method for implementing an end-to-end automatic speech translation (AST) model with a neural transducer, the method comprising:
accessing a training dataset comprising an audio dataset comprising spoken language utterances in a first language and a text dataset comprising transcription labels in a second language, the transcription labels corresponding to the spoken language utterances;
access an automatic speech translation (AST) model that is configured to receive and encode audio data, the AST model being further configured to generate translated language based on transcription label output;
accessing an end-to-end AST model based on a neural transducer comprising at least an acoustic encoder which is configured to receive and encode audio data, a prediction network which is integrated in a parallel model architecture with the acoustic encoder in the end-to-end AST model and configured to predict a subsequent language token based on a previous transcription label output;
train the AST model by applying the training dataset to the AST model;
applying the training dataset to the end-to-end AST model;
access input audio data that is in the first language; and
use the AST model to generate a transcription of the input audio data, the transcription of the input audio data being in the second language.
generating a transcription in the second language of input audio data in the first language based on the trained end-to-end AST model.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 – 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
Independent claims 1, 9 and 17 recite the limitations of accessing training dataset that comprises audio of spoken language in a first language as well as a transcription label of the audio in a second language, accessing an automatic speech translation model which has been configured to receive and encode audio data and also generate translated language based on a transcription label output, training of the automatic speech translation model by applying the mentioned training dataset, accessing input audio data in a first language, and then applying the automatic speech translation model to generate a transcription of the input audio data in a second language.
Nothing in the claims precludes the claims from being performed in the human mind. The entire process involves data collection, data retrieval, and data generation. A human may obtain spoken audio information in a first language and text data matched to the spoken audio but presented in a second language, the human may then listen to and read through several audio and text data that match text data in the second language to audio data in the first language in order to gain knowledge about the translation process, such that the spoken audio data is received and changed to a format which the human can properly understand, the human learns how the translation works, the human then receives input audio data in the first language and based on the information which the human has learnt, generates a transcription of the input audio data in the second language. The claims hereby recite a mental process.
This judicial exception is not integrated into a practical application as the claims simply teach of collecting data in the form of collecting audio dataset in a first language, text dataset in a second language, and also the training of an automatic speech translation model; retrieving data through accessing an automatic speech translation model and accessing input audio; and generating data by generating a transcription in a second language of audio data in a first language, based on accessing an automatic speech translation model. The claims make mention of an automatic speech translation model but this is presented in a generic form.
The invention is not tied to any particular defining structure and simply provides instructions to apply the judicial exception. The technique can be performed by a generic computer which would be presented as a tool to implement the abstract idea (classifiable as automation of the mental process steps). The Specification in [00100] provides a general-purpose computer that includes computer hardware, which can be a generic computer. The claims also refer to an automatic speech translation model which is based on a neural transducer [0006] and also based on a recurrent neural network transducer and a transformer transducer [0037]. The automatic speech translation model is provided in generic terms such that by its presentation here in the claims, it can be performed by a human mind. Mentioning the automatic speech translation model can simply refer to the application of a general-purpose computer to perform a mental process of translating audio speech data in a first language into text data in a second language, which is a mental process. A human may perform this by receiving the speech data in one language, understanding what’s being said, and converting it by writing out the text in a second language. The one or more processors, hardware storage devices and the automatic speech translation model are recited at a high level of generality that they amount to no more than mere instructions to apply the exception using a generic computer. The claims do not include any additional element that would that would be sufficient to amount to significantly more than the judicial exception because the invention is not tied to a practical application.
The claims provide techniques that amount to no more than mere instructions that apply the judicial exception which can be performed by a generic device. Merely mentioning the processing circuitry and memory amounts to no more than general-purpose hardware used as tools to implement the abstract idea and does not provide any particular application other than applying it for the purpose of implementing a judicial exception. Mere instructions to apply an exception using a generic device cannot provide an inventive concept. Claims 1, 9 and 17 are not eligible.
Claims 2, 10 and 18 provide that the AST model is an end-to-end speech translation model based on an RNN transducer structure and/or a transformer-transducer structure. This simply provides the generic recitation of an RNN transducer or transformer-transducer to be applied to an AST model. These models are mentioned in [0031] and [0037] of the Specification without any particular information on the way they are being trained or the way they are being applied to performing automatic speech translation. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claims 3, 11 and 19 provide that the speech translation model requires no attention and leads to handling word reordering during translation. A human may properly handle word reordering during translation. That the model requires no attention indicates that it is trained without an attention mechanism that would let it weigh different parts of the input sequence when producing its output. This is also a process which a human can take into consideration when mentally learning the translation, resulting in a process which a human can perform without an attention mechanism. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claims 4, 12 and 20 provide that the end-to-end speech translation model is designed for natural streaming. A human may perform a live translation of received audio data. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claims 5 and 13 provide that the end-to-end speech translation model uses augmented data for training instead of raw speech translation data. A human may, over time, prepare augmented data that adds nuances and different linguistic patterns and phrasing, to be used while learning speech translation. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claims 6 and 14 provide that the end-to-end speech translation model uses synthesised data for training. A human may obtain synthesised data that adds different information which may take a long time to naturally obtain or prepare, to be used while learning speech translation. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claims 7 and 15 provide that the end-to-end speech translation model is based on the transformer-transducer structure. This simply provides the generic recitation of a transformer-transducer to be applied to an AST model. This model is mentioned in [0037] of the Specification without any particular information on the way it is being trained or the way they are being applied to performing automatic speech translation. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claims 8 and 16 provide that the end-to-end speech translation model is based on the recurrent neural network transducer structure. This simply provides the generic recitation of a transformer-transducer to be applied to an AST model. This model is mentioned in [0031] and [0037] of the Specification without any particular information on the way it is being trained or the way they are being applied to performing automatic speech translation. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 9 and 17 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by INDURTHI et al. (US 2021/0065690 A1: hereafter — Indurthi).
For claim 1, discloses a method comprising:
accessing a training dataset comprising (i) an audio dataset comprising a spoken language utterance spoken in a first language and (ii) a text dataset comprising a transcription label that is written in a second language, the transcription label corresponding to the spoken language utterance (Indurthi: [0004]; [0019] — trained speech model being updated based on text in a second language which is a converted form of speech in a first language as well as text in the second language, also comprising speech in a first language (as provided by its first information); [0075] — training a speech translation model by making use of speech in a first language and text in a second language, the text in the second language corresponding to the speech in the first language);
accessing an automatic speech translation (AST) model that is configured to receive and encode audio data, the AST model being further configured to generate translated language based on transcription label output (Indurthi: [0072] — an encoder which receives input speech (the input speech is in the first language); [0023] — the speech translation model is trained to convert speech in a first language into a text in a second language, and output the text; [0080] — describes an automatic speech translation model which starts with automatic speech recognition, and then machine translation);
training the AST model by applying the training dataset to the AST model (Indurthi: [0019] — updating the training of speech translation model based on text in a second language which is a converted form of speech in a first language as well as text in the second language, also comprising speech in a first language (as provided by its first information));
accessing input audio data that is in the first language (Indurthi: [0075] — receiving user speech input in a first language); and
using the AST model to generate a transcription of the input audio data, the transcription of the input audio data being in the second language (Indurthi: [0072] — the output being text that are a translation of the speech input; [0110] — ‘the speech translation model may be trained to convert a speech in a first language into a text in a second language and output the text’).
As for claim 9, computer system claim 9 and method claim 1 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Indurthi in [0009] provides both a processor and storage memory suitable to read upon the limitations of this claim. Accordingly, claim 9 is similarly rejected under the same rationale as applied above with respect to method claim 1.
As for claim 17, computer program product claim 17 and method claim 1 are related as computer program product storing executable instructions required for performing the claimed method steps on a computer. Indurthi in [0009] provides both a processor and storage memory, and a non-transitory storage medium in [0122], suitable to read upon the limitations of this claim Accordingly, claim 17 is similarly rejected under the same rationale as applied above with respect to method claim 1.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 4, 8, 10, 12, 16 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Indurthi (US 2021/0065690 A1) as applied to claim 1, in view of Han et al. (US 2020/0126538 A1: hereafter — Han).
For claim 2, claim 1 is incorporated and the reference of Indurthi discloses an automatic speech translation process that is trained by using a sequence-to-sequence model. This reference however differs from the claimed invention in that the claimed invention further provides teaching for an end-to-end speech translation model that is based on a recurrent neural network transducer structure.
This is however not new to the art as the reference of Han is now introduced to teach this as:
the method, wherein the AST model is an end-to-end speech translation model that is based on a recurrent neural network transducer structure and/or a transformer-transducer structure (Han: [0025] — sequence-to-sequence models (that can be applied to speech translation) including a recurrent neural network transducer).
Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of Indurthi which provides an automatic speech translation process trained by using a sequence-to-sequence model, by applying the known technique of Han which bases the automatic speech translation model on a recurrent neural network transducer structure, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of enabling an end-to-end streaming speech translation that reduces latency and aids in avoiding speech recognition error propagation. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007).
For claim 4, claim 2 is incorporated and the combination of Indurthi in view of Han discloses the method, wherein the end-to-end speech translation model is designed for natural streaming (Han: [0045] — a streaming attention-based model, implemented by a neural transducer).
For claim 8, claim 2 is incorporated and the combination of Indurthi in view of Han discloses the method, wherein the end-to-end speech translation model is based on the recurrent neural network transducer structure (Han: [0025] — sequence-to-sequence models (that can be applied to speech translation) including a recurrent neural network transducer).
As for claim 10, computer system claim 10 and method claim 2 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 10 is similarly rejected under the same rationale as applied above with respect to method claim 2.
As for claim 12, computer system claim 12 and method claim 4 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 12 is similarly rejected under the same rationale as applied above with respect to method claim 4.
As for claim 16, computer system claim 16 and method claim 8 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 16 is similarly rejected under the same rationale as applied above with respect to method claim 8.
As for claim 18, computer program product claim 18 and method claim 2 are related as computer program product storing executable instructions required for performing the claimed method steps on a computer. Accordingly, claim 18 is similarly rejected under the same rationale as applied above with respect to method claim 2.
As for claim 20, computer program product claim 20 and method claim 4 are related as computer program product storing executable instructions required for performing the claimed method steps on a computer. Accordingly, claim 20 is similarly rejected under the same rationale as applied above with respect to method claim 4.
Claims 3, 11 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Indurthi (US 2021/0065690 A1) in view of Han (US 2020/0126538 A1) as applied to claim 2, further in view of Chuang, Shun-Po, et al. (“Investigating the reordering capability in CTC-based non-autoregressive end-to-end speech translation.” Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2021: hereafter — Chuang).
For claim 3, claim 2 is incorporated and the combination of Indurthi in view of Han provides teaching for end-to-end speech translation. This combination however differs from the claimed invention in that the claimed invention further provides teaching for a speech translation model that requires no attention and preserves a full capacity of the end-to-end speech translation model to handle word reordering during translation.
This is however not new to the art as the reference of Chuang is now introduced to teach this as:
the method, wherein the end-to-end speech translation model requires no attention, resulting in preservation of a full capability of the end-to-end speech translation model to handle word reordering during translation (Chuang: Page 1068 Col 2 Par 1 — ‘In this work, we use connectionist temporal classification (CTC) (Graves et al., 2006) to train NAR models for ST’ (noting that CTCs do not require n attention mechanism to function); Page 1068 Col 2 Par 3 — the use of CTC leading to achieving non-monotonic mapping (the mapping being one between the input sequence and the output sequence); Page 1070 Col 1 Par 1 — a model that handles reordering).
Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Indurthi in view of Han which provides an end-to-end speech translation, by applying the known technique of Chuang which provides a speech translation model that requires no attention and preserves a full capacity of the end-to-end speech translation model to handle word reordering during translation, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of handling real-time (streaming) speech translation applications. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007).
As for claim 11, computer system claim 11 and method claim 3 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 11 is similarly rejected under the same rationale as applied above with respect to method claim 3.
As for claim 19, computer program product claim 19 and method claim 3 are related as computer program product storing executable instructions required for performing the claimed method steps on a computer. Accordingly, claim 19 is similarly rejected under the same rationale as applied above with respect to method claim 3.
Claims 5, 6, 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Indurthi (US 2021/0065690 A1) in view of Han (US 2020/0126538 A1) as applied to claim 2, further in view of Di Gangi, Mattia A., et al. (“Data augmentation for end-to-end speech translation: FBK@ IWSLT ‘19.” Proceedings of the 16th International Conference on Spoken Language Translation. 2019: hereafter — Di Gangi).
For claim 5, claim 2 is incorporated and the combination of Indurthi in view of Han provides teaching for an end-to-end speech translation trained with audio dataset spoken in a first language, and text of the audio, in a second language. This however differs from the claimed invention in that the claimed invention further provides teaching for performing the training using augmented data.
This is however not new to the art as the reference of Di Gangi is now introduced to teach this as:
the method, wherein the end-to-end speech translation model uses only augmented data for training instead of using raw speech translation data (Di Gangi: Abstract — ‘Finally, we trained our models using SpecAugment, an augmentation technique that randomly masks portions of the spectrograms in order to make them different at every training epoch’; 2.3 — making use of SpecAugment which is a data augmentation technique; 5.2 — training with SpecAugment).
Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Indurthi in view of Han which provides an end-to-end speech translation trained with audio dataset spoken in a first language, and text of the audio, in a second language, by applying the known technique of Di Gangi which presents the use of augmented data to be applied to training an end-to-end speech translation model, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of increasing the type of data being made available for training by providing data from different domains, without needing to obtain possibly costly new/original datasets. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007).
For claim 6, claim 2 is incorporated and the combination of Indurthi in view of Han provides teaching for an end-to-end speech translation trained with audio dataset spoken in a first language, and text of the audio, in a second language. This however differs from the claimed invention in that the claimed invention further provides teaching for performing the training using synthesised data.
This is however not new to the art as the reference of Di Gangi is now introduced to teach this as:
the method, wherein the end-to-end speech translation model uses synthesized data for training (Di Gangi: Abstract — training using synthetic corpus; 2.1 — generating synthetic data to be applied to training, an end-to-end speech translation model that was trained using synthetic data; 5.1 — training based on synthetic data).
Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Indurthi in view of Han which provides an end-to-end speech translation trained with audio dataset spoken in a first language, and text of the audio, in a second language, by applying the known technique of Di Gangi which presents the use of synthetic data to be applied to training an end-to-end speech translation model, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of increasing the type of data being made available for training by providing data from different domains, without needing to obtain possibly costly new/original datasets. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007).
As for claim 13, computer system claim 13 and method claim 5 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 13 is similarly rejected under the same rationale as applied above with respect to method claim 5.
As for claim 14, computer system claim 14 and method claim 6 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 14 is similarly rejected under the same rationale as applied above with respect to method claim 6.
Claims 7 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Indurthi (US 2021/0065690 A1) in view of Han (US 2020/0126538 A1) as applied to claim 2, further in view of WU et al. (US 2022/0351718 A1: hereafter — Wu).
For claim 7, claim 1 is incorporated and the combination of Indurthi in view of Han provides teaching for an end-to-end speech translation model based on a recurrent neural network transducer structure. This combination however differs from the claimed invention in that the claimed invention further provides teaching for the translation being based on a transformer transducer structure.
This is however not new to the art, as the reference of Wu is now introduced to teach this as:
the method, wherein the end-to-end speech translation model is based on the transformer-transducer structure (Wu: [0013] — a transformer-transducer-based deep neural network architecture for training an E2E ASR model).
Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to combine the teaching of the combination of Indurthi in view of Han which bases the automatic end-to-end speech translation model on a recurrent neural network transducer structure, with the known teaching of Wu which provides the end-to-end speech translation model based on a transformer-transducer structure, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of providing direct mapping of raw audio data to a target language text while avoiding errors that could result from speech recognition and machine translations, while also reducing translation latency. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007).
As for claim 15, computer system claim 15 and method claim 7 are related as system and the method of using same, with each claimed element’s function corresponding to the claimed method step. Accordingly, claim 15 is similarly rejected under the same rationale as applied above with respect to method claim 7.
Conclusion
The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure.
MATUSOV et al. (US 2020/0226327 A1) provides an end-to-end multi-lingual speech-to-speech system [0008] having a system that uses real sentence triples in the form of source language speech, its transcript, its translation [0054].
Dhawan et al. (US 2022/0115028 A1) provides teaching for a speech to speech translation system which requires parallel audio samples in two languages, the training involving an equal augmentation of both audio data (Abstract).
Slawson et al. (US 2009/0240539 A1) provides teaching for a speech-to-text translation which augments its training with additional demographic information [0033].
Any inquiry concerning this communication or earlier communications from the Examiner should be directed to OLUWADAMILOLA M. OGUNBIYI whose telephone number is (571)272-4708. The Examiner can normally be reached Monday – Thursday (8:00 AM – 5:30 PM Eastern Standard Time).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s Supervisor, PARAS D. SHAH can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OLUWADAMILOLA M OGUNBIYI/Examiner, Art Unit 2653