DETAILED ACTION
This action is in response to the initial filing of application no. 19/100,543 on 01/31/2025.
Claims 1-7, 9, 10 and 12 -22 are still pending in this application, with claims 1, 9 and 10 being independent.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 4, 7,14, 17 and 20 are objected to as being dependent upon rejected base claims, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim 4 recites, wherein determining the undetermined image sample feature, the first undetermined source text, and the first undetermined target text according to the sample data comprises: when the sample data comprises the multimodal multilingual data, using an image feature output after the first image being input into the image encoder as the undetermined image sample feature, using the first source-language text as the first undetermined source text, and using the first target-language text as the first undetermined target text; or when the sample data comprises the unimodal multilingual data, using a preset image sample feature as the undetermined image sample feature, using the second source-language text as the first undetermined source text, and using the second target-language text as the first undetermined target text; or when the sample data comprises the multimodal monolingual data, using an image feature output after the second image being input into the image encoder as the undetermined image sample feature, using a text obtained by masking the third target-language text as the first undetermined source text, and using the third target-language text as the first undetermined target text.
In particular, the prior art fails to teach the following: when the sample data comprises the multimodal monolingual data, using a text obtained by masking the third target-language text as the first undetermined source text. (Kwon teaches the multimodal multilingual data limitations,4.1 Data and 4.3 Experimental Settings) (Gronroos teaches the unimodal multilingual data limitations using a dummy visual feature as the preset image sample feature, 4.3 Training, pg. 608 and 609)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 9, 10, 12 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kwon et al. (“A text-based visual context modulation neural model for multimodal machine translation”) (“Kwon”) in view of Gronroos et al. (The MeMAD Submission to the WMT18 Multimodal Translation Task) (“Gronroos”) and further in view of Meng et al. (US 2022/0245365) (“Meng”).
For claim 1, Kwon discloses a translation method (Abstract), comprising: a pre-generated target multimodal translation model (target multiplied Transformer with modulation network, Fig.1 and Fig.2b; 3.1.2 Transformer, 3.2 Text-based modulating network, 4.2 Baseline, pg. 214 and 215) which accepts an input source test and source associated image and obtains a target translation text output (4.1 Data ,4.3 Experimental Setting, 4.4 Result and Table 2, pg. 215 - 217), wherein the pre-generated target multimodal translation model is a model generated by training an undetermined multimodal translation model according to sample data (The model is trained, validated and tested using in-domain and out-domain data., 3.2 Text based modulating network, 4.1 Data, 4.3 Experimental setting, pg. 214 - 216), the sample data comprises at least two of multimodal multilingual data (Multi30k, 4.1 Data) (Multi30k, 4.1 Data), the multimodal multilingual data comprises a first source-language text, a first target-language text, and a first image corresponding to the first source-language text (The triplet of Multi30k dataset, 4.1 Data), the first image is an associated image corresponding to the first source-language text (Multi30k dataset, 4.1 Data).
Yet, Kwon fails to teach the following: determining a source text to be translated and a source associated image corresponding to the source text; and inputting the source text and the source associated image into the pre-generated target multimodal translation model; the sample data further comprises unimodal multilingual data, and multimodal monolingual data, the unimodal multilingual data comprises a second source-language text and a second target- language text, the multimodal monolingual data comprises a third target-language text and a second image, the second image is an associated image corresponding to the third target-language text, and a language type of the first source-language text, a language type of the second source-language text, and a language type of the source text are the same, and a language type of the first target- language text, a language type of the second target-language text, a language type of the third target-language text, and a language type of the target translation text are the same.
However, Gronroos discloses a multimodal machine translation system and method (Abstract), comprising the following: a multimodal machine translation module (1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.1 Architecture, 4.2 Visual feature selection, pg. 603,604, 606 - 609) is generated using multimodal multilingual data (Multi30K, 1. Introduction, pg. 603), multimodal monolingual data (MS-COCO, 1. Introduction) and unimodal multilingual data (OpenSubtitles, 1.Introduction) (4.3 Training and 4.4 Results, pg. 603, 608 and 609); the multimodal monolingual data comprises a target-language text and an image (1. Introduction, pg. 603), wherein the image is an associated image corresponding to the target-language text (1.Introduction, pg. 603); the unimodal multilingual data comprises a source-language text and a second target- language text (“Our initial experiments showed that movie subtitles and their translations work rather well to augment the given training data. Therefore, we included parallel subtitles from the OpenSubtitles2018 corpus.”, 1. Introduction and 2. Experiment, pg. 603 and 604); a language type (German/French) of the source language text of the multimodal multilingual data and a language type of the source language text of the unimodal multilingual data (Table 1, 1. Introduction, pg. 603); and a language type of a target- language text of the multimodal multilingual data, a language type of the target-language text of the unimodal multilingual data, and a language type (English) of the target-language text of the multimodal monolingual data is the same (Table 1, 1. Introduction, pg. 603).
Moreover, Meng discloses a translation method and apparatus (Abstract), comprising a generated target multimodal machine translation model (Fig.1; [0150]) which performs the following: determine a source text to be translated and a source associated image corresponding the to the source text ([0079] [0081]); inputting the source text and the associated source image into the generated target multimodal translation model ([0088 – 0106]) to obtain a target translation text output by the target multimodal translation model ([0106 – 0111]), wherein a language type of the source text and target text is the same as the language types of data used to trained the model ([0004] [0172] [0174]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve Kwon’s invention in the same way that Gronroos’s invention has been improved to achieve the following, predictable results for the purpose of improving the performance of the multimodal machine translation model by augmenting the triplet dataset used to train the model (Kwon, 1.Introduction) (Gronroos, 2. Experiment 1: Optimizing the Text-Based Machine Translation, pg. 603 – 605): the sample data further comprises unimodal multilingual data, and multimodal monolingual data, the unimodal multilingual data comprises a second source-language text and a second target- language text, the multimodal monolingual data comprises a third target-language text and a second image, the second image is an associated image corresponding to the third target-language text, and a language type of the first source-language text and a language type of the second source-language text are the same, and a language type of the first target- language text, a language type of the second target-language text and a language type of the third target-language text are the same.
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kwon and Gronroos in the same way that the Meng’s invention has been improved to achieve the following, predictable results for the purpose of performing multimodal machine translation with improved results by using an architecture that modulates image vectors conditioned on their paired text information (Kwon, 1. Introduction): further determining a source text to be translated and a source associated image corresponding to the source text; and further inputting the source text and the source associated image into the pre-generated target multimodal translation model, wherein a language type of the source text is the same as the first source-language text (training data) and a language type of the target translation text is the same as the first target- language text (training data).
For claim 9, Kwon discloses a translation method (Abstract), comprising: a pre-generated target multimodal translation model (target multiplied Transformer with modulation network, Fig.1 and Fig.2b; 3.1.2 Transformer, 3.2 Text-based modulating network, 4.2 Baseline, pg. 214 and 215) which accepts an input source test and source associated image and obtains a target translation text output (4.1 Data ,4.3 Experimental Setting, 4.4 Result and Table 2, pg. 215 - 217), wherein the pre-generated target multimodal translation model is a model generated by training an undetermined multimodal translation model according to sample data (The model is trained, validated and tested using in-domain and out-domain data., 3.2 Text based modulating network, 4.1 Data, 4.3 Experimental setting, pg. 214 - 216), the sample data comprises at least two of multimodal multilingual data (Multi30k, 4.1 Data) (Multi30k, 4.1 Data), the multimodal multilingual data comprises a first source-language text, a first target-language text, and a first image corresponding to the first source-language text (The triplet of Multi30k dataset, 4.1 Data), the first image is an associated image corresponding to the first source-language text (Multi30k dataset, 4.1 Data).
Yet, Kwon fails to teach the following: a non-transitory computer readable medium having a computer program stored thereon, wherein the computer program, when executed by a processing apparatus, implements the translation method; determining a source text to be translated and a source associated image corresponding to the source text; and inputting the source text and the source associated image into the pre-generated target multimodal translation model; the sample data further comprises unimodal multilingual data, and multimodal monolingual data, the unimodal multilingual data comprises a second source-language text and a second target- language text, the multimodal monolingual data comprises a third target-language text and a second image, the second image is an associated image corresponding to the third target-language text, and a language type of the first source-language text, a language type of the second source-language text, and a language type of the source text are the same, and a language type of the first target- language text, a language type of the second target-language text, a language type of the third target-language text, and a language type of the target translation text are the same.
However, Gronroos discloses a multimodal machine translation system and method (Abstract), comprising the following: a multimodal machine translation module (1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.1 Architecture, 4.2 Visual feature selection, pg. 603,604, 606 - 609) is generated using multimodal multilingual data (Multi30K, 1. Introduction, pg. 603), multimodal monolingual data (MS-COCO, 1. Introduction) and unimodal multilingual data (OpenSubtitles, 1.Introduction) (4.3 Training and 4.4 Results, pg. 603, 608 and 609); the multimodal monolingual data comprises a target-language text and an image (1. Introduction, pg. 603), wherein the image is an associated image corresponding to the target-language text (1.Introduction, pg. 603); the unimodal multilingual data comprises a source-language text and a second target- language text (“Our initial experiments showed that movie subtitles and their translations work rather well to augment the given training data. Therefore, we included parallel subtitles from the OpenSubtitles2018 corpus.”, 1. Introduction and 2. Experiment, pg. 603 and 604); a language type (German/French) of the source language text of the multimodal multilingual data and a language type of the source language text of the unimodal multilingual data (Table 1, 1. Introduction, pg. 603); and a language type of a target- language text of the multimodal multilingual data, a language type of the target-language text of the unimodal multilingual data, and a language type (English) of the target-language text of the multimodal monolingual data is the same (Table 1, 1. Introduction, pg. 603).
Moreover, Meng discloses a translation method and apparatus (Abstract), comprising a generated target multimodal machine translation model (Fig.1; [0150]) which performs the following translation method: determine a source text to be translated and a source associated image corresponding the to the source text ([0079] [0081]); inputting the source text and the associated source image into the generated target multimodal translation model ([0088 – 0106]) to obtain a target translation text output by the target multimodal translation model ([0106 – 0111]), wherein a language type of the source text and target text is the same as the language types of data used to trained the model ([0004] [0172] [0174]). Additionally, Meng discloses a non-transitory computer readable medium having a computer program stored thereon, wherein the computer program, when executed by a processing apparatus, implements the translation method ([0020]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve Kwon’s invention in the same way that Gronroos’s invention has been improved to achieve the following, predictable results for the purpose of improving the performance of the multimodal machine translation model by augmenting the triplet dataset used to train the model (Kwon, 1.Introduction) (Gronroos, 2. Experiment 1: Optimizing the Text-Based Machine Translation, pg. 603 – 605) : the sample data further comprises unimodal multilingual data, and multimodal monolingual data, the unimodal multilingual data comprises a second source-language text and a second target- language text, the multimodal monolingual data comprises a third target-language text and a second image, the second image is an associated image corresponding to the third target-language text, and a language type of the first source-language text and a language type of the second source-language text are the same, and a language type of the first target- language text, a language type of the second target-language text and a language type of the third target-language text are the same.
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kwon and Gronroos in the same way that the Meng’s invention has been improved to achieve the following, predictable results for the purpose of performing multimodal machine translation with improved results by using an architecture that modulates image vectors conditioned on their paired text information (Kwon, 1. Introduction): further determining a source text to be translated and a source associated image corresponding to the source text; and further inputting the source text and the source associated image into the pre-generated target multimodal translation model, wherein a language type of the source text is the same as the first source-language text (training data) and a language type of the target translation text is the same as the first target- language text (training data), wherein the translation method is implemented by a processing apparatus which executes a non-transitory computer readable medium having a computer program stored thereon.
For claim 10, Kwon discloses a translation method (Abstract), comprising: a pre-generated target multimodal translation model (target multiplied Transformer with modulation network, Fig.1 and Fig.2b; 3.1.2 Transformer, 3.2 Text-based modulating network, 4.2 Baseline, pg. 214 and 215) which accepts an input source test and source associated image and obtains a target translation text output (4.1 Data ,4.3 Experimental Setting, 4.4 Result and Table 2, pg. 215 - 217), wherein the pre-generated target multimodal translation model is a model generated by training an undetermined multimodal translation model according to sample data (The model is trained, validated and tested using in-domain and out-domain data., 3.2 Text based modulating network, 4.1 Data, 4.3 Experimental setting, pg. 214 - 216), the sample data comprises at least two of multimodal multilingual data (Multi30k, 4.1 Data) (Multi30k, 4.1 Data), the multimodal multilingual data comprises a first source-language text, a first target-language text, and a first image corresponding to the first source-language text (The triplet of Multi30k dataset, 4.1 Data), the first image is an associated image corresponding to the first source-language text (Multi30k dataset, 4.1 Data).
Yet, Kwon fails to teach the following: an electronic device comprising a memory having a computer program stored thereon, and a processing apparatus configured to execute the computer program in the memory to implement the translation method; determining a source text to be translated and a source associated image corresponding to the source text; and inputting the source text and the source associated image into the pre-generated target multimodal translation model; the sample data further comprises unimodal multilingual data, and multimodal monolingual data, the unimodal multilingual data comprises a second source-language text and a second target- language text, the multimodal monolingual data comprises a third target-language text and a second image, the second image is an associated image corresponding to the third target-language text, and a language type of the first source-language text, a language type of the second source-language text, and a language type of the source text are the same, and a language type of the first target- language text, a language type of the second target-language text, a language type of the third target-language text, and a language type of the target translation text are the same.
However, Gronroos discloses a multimodal machine translation system and method (Abstract), comprising the following: a multimodal machine translation module (1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.1 Architecture, 4.2 Visual feature selection, pg. 603,604, 606 - 609) is generated using multimodal multilingual data (Multi30K, 1. Introduction, pg. 603), multimodal monolingual data (MS-COCO, 1. Introduction) and unimodal multilingual data (OpenSubtitles, 1.Introduction) (4.3 Training and 4.4 Results, pg. 603, 608 and 609); the multimodal monolingual data comprises a target-language text and an image (1. Introduction, pg. 603), wherein the image is an associated image corresponding to the target-language text (1.Introduction, pg. 603); the unimodal multilingual data comprises a source-language text and a second target- language text (“Our initial experiments showed that movie subtitles and their translations work rather well to augment the given training data. Therefore, we included parallel subtitles from the OpenSubtitles2018 corpus.”, 1. Introduction and 2. Experiment, pg. 603 and 604); a language type (German/French) of the source language text of the multimodal multilingual data and a language type of the source language text of the unimodal multilingual data (Table 1, 1. Introduction, pg. 603); and a language type of a target- language text of the multimodal multilingual data, a language type of the target-language text of the unimodal multilingual data, and a language type (English) of the target-language text of the multimodal monolingual data is the same (Table 1, 1. Introduction, pg. 603).
Moreover, Meng discloses a translation method and apparatus (Abstract), comprising a generated target multimodal machine translation model (Fig.1; [0150]) which performs the following: determine a source text to be translated and a source associated image corresponding the to the source text ([0079] [0081]); inputting the source text and the associated source image into the generated target multimodal translation model ([0088 – 0106]) to obtain a target translation text output by the target multimodal translation model ([0106 – 0111]), wherein a language type of the source text and target text is the same as the language types of data used to trained the model ([0004] [0172] [0174]). Additionally, Meng discloses a non-transitory computer readable medium having a computer program stored thereon, wherein the computer program, when executed by a processing apparatus, implements the translation method ([0020]). Additionally, Meng discloses an electronic device comprising a memory having a computer program stored thereon, and a processing apparatus configured to execute the computer program in the memory to implement the translation method ([0016 – 0019]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve Kwon’s invention in the same way that Gronroos’s invention has been improved to achieve the following, predictable results for the purpose of improving the performance of the multimodal machine translation model by augmenting the triplet dataset used to train the model (Kwon, 1.Introduction) (Gronroos, 2. Experiment 1: Optimizing the Text-Based Machine Translation, pg. 603 – 605) : the sample data further comprises unimodal multilingual data, and multimodal monolingual data, the unimodal multilingual data comprises a second source-language text and a second target- language text, the multimodal monolingual data comprises a third target-language text and a second image, the second image is an associated image corresponding to the third target-language text, and a language type of the first source-language text and a language type of the second source-language text are the same, and a language type of the first target- language text, a language type of the second target-language text and a language type of the third target-language text are the same.
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kwon and Gronroos in the same way that the Meng’s invention has been improved to achieve the following, predictable results for the purpose of performing multimodal machine translation with improved results by using an architecture that modulates image vectors conditioned on their paired text information (Kwon, 1. Introduction): further determining a source text to be translated and a source associated image corresponding to the source text; and further inputting the source text and the source associated image into the pre-generated target multimodal translation model, wherein a language type of the source text is the same as the first source-language text (training data) and a language type of the target translation text is the same as the first target- language text (training data), wherein the translation method is implemented electronic device comprising a memory having a computer program stored thereon and a processing apparatus configured to execute the computer program in the memory.
For claims 2, 12, and 18, Kwon and Meng further disclose, wherein the target multimodal translation model comprises an image encoder (Kwon, CNN ResNet-152 in a trg-mul transformer network - based machine translation model with a modulation network, Fig.1 and Fig.2b; 2.1 Multimodal machine translation, 3.2 Text-based modulation network, 4.1 Data, 4.2 Baseline), a first text encoder (Kwon, source encoder in a trg-mul transformer network- based machine translation model with a modulation network, Fig.1 and Fig.2b; 2.1 Multimodal machine translation, 3.1.2. Transformer, 3.2 Text-based modulation network, 4.1 Data, 4.2 Baseline), and a first text decoder (Kwon, target decoder in a trg-mul transformer network-based machine translation model with a modulation network, Fig.1 and Fig.2b; 2.1 Multimodal machine translation, 3.1.2. Transformer , 3.2 Text-based modulation network, 4.1 Data, 4.2 Baseline) and inputting the source text and the source associated image into the pre-generated target multimodal translation model to obtain the target translation text output by the target multimodal translation model comprises: inputting the source associated image (Kwon, Fig.2b, Image) into the image encoder to obtain a source associated image feature output by the image encoder (Kwon, Fig.1, Fig.2b; 3.2 Text-based modulating network) (Meng, [0079] [0081]); inputting the source text (Kwon, Fig.2b, source sentence) into the first text encoder to obtain a first source text feature output by the first text encoder (Kwon, Fig.1, Fig.2b; 3.2 Text-based modulating network) (Meng, [0079] [0081]); obtaining a target feature by weighting the first source text feature and the source associated image feature (Kwon, Fig.1, Fig.2b; 3.2 Text-based modulating network); and inputting the target feature into the first text decoder to obtain the target translation text output by the first text decoder (Kwon, The target feature is used to modulate a target sentence embedding which is input to the text decoder. Thus, the target feature is input to the text decoder as the modulation values., Fig.2B; 4.2 Baseline,4.3 Experimental setting) .
Claim(s) 3, 13 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over in view of Kwon et al. (“A text-based visual context modulation neural model for multimodal machine translation”) (“Kwon”) in view of Gronroos et al. (The MeMAD Submission to the WMT18 Multimodal Translation Task) (“Gronroos”), and further in view of Meng et al. (US 2022/0245365) (“Meng”) and further in view of Hu et al. (US 2023/0120631) (“Hu”).
For claims 3, 13 and 19, the combination of Kwon, Gronroos and Meng further discloses, wherein the undetermined multimodal translation model comprises the image encoder, the first text encoder, and the first text decoder, and the target multimodal translation model is generated by: obtaining the sample data (Kwon, 4.1 Data, 4.3 Experimental Setting) (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609); and cyclically performing a first training step according to the sample data (Kwon, 4.3 Experimental setting) (Gronroos, 4.3 Training, pg. 608 and 609), and using the trained undetermined multimodal translation model as the target multimodal translation model (Kwon, Fig.2b) (Meng, Fig.1; [0040] [0150] [0186]), wherein the first training step comprises: determining an undetermined image sample feature, a first undetermined source text , and a first undetermined target text according to the sample data (Kwon, 4.1 Data, 4.3 Experimental Setting) (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609); inputting the first undetermined source text into the first text encoder to obtain a first text sample feature output by the first text encoder (Kwon, Fig.2b; 4.1 Data, 4.3 Experimental Setting) (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609); determining a first multimodal sample feature according to the first text sample feature and the first image sample feature (Kwon, Fig.2b; 3.2 Text-based modulating network, 4.1 Data, 4.3 Experimental Setting); inputting the first multimodal sample feature into the first text decoder to obtain a first translation text output by the first text decoder (Kwon, Fig.2b; 3.2 Text-based modulating network, 4.1 Data, 4.3 Experimental Setting)
Yet, the combination of Kwon, Gronroos and Meng fails to teach the following: cyclically performing the first training step until it is determined that the trained undetermined multimodal translation model meets a first preset stop condition;
determining a first loss value according to the first translation text and the first undetermined target text, and when it is determined, according to the first loss value, that the undetermined multimodal translation model does not meet the first preset stop iteration condition, updating a parameter of the undetermined multimodal translation model according to the first loss value to obtain a trained undetermined multimodal translation model, and using the trained undetermined multimodal translation model as a new undetermined multimodal translation model.
However, Hu discloses a neural network training method (Abstract), comprising the following: determining a first loss value according to a result label value of training data and a result output by a neural network model ([0172]), and when it is determined, according to the first loss value, that the undetermined neural network model does not meet the first preset stop iteration condition, updating a parameter of the undetermined neural network model according to the first loss value to obtain a trained undetermined neural network model ([0172] [0193] [0194] [0243 – 0247]) and using the trained undetermined multimodal translation model as a new undetermined multimodal translation model ([0239 – 0242] [0248]). Furthermore, Hu discloses cyclically performing the first training step until it is determined that the trained undetermined multimodal translation model meets a first preset stop condition ([0243 - 0248]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kwon, Gronroos and Meng in the same way that Hu’s invention has been improved to achieve the following, predictable results for the purpose of efficiently training the undetermined multimodal translation model to achieve a desired translation quality: cyclically performing the first training step until it is determined that the trained undetermined multimodal translation model meets a first preset stop condition; determining a first loss value according to an output, e.g. first translation text, of the undetermined multimodal translation model and a result label value, e.g. first undetermined target text, of training data, and when it is determined, according to the first loss value, that the undetermined multimodal translation model does not meet a first preset stop iteration condition, updating a parameter of the undetermined multimodal translation model according to the first loss value to obtain a trained undetermined multimodal translation model, and using the trained undetermined multimodal translation model as a new undetermined multimodal translation model.
Claim(s) 5, 15 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kwon et al. (“A text-based visual context modulation neural model for multimodal machine translation”) (“Kwon”) in view of Gronroos et al. (The MeMAD Submission to the WMT18 Multimodal Translation Task) (“Gronroos”), and further in view of Meng et al. (US 2022/0245365) (“Meng”) and further in view of Parida et al. (“Multimodal Neural Machine Translation System for English to Bengali”).
For claims 5, 15 and 21, the combination of Kwon, Gronroos and Meng fails to teach the following: wherein the target multimodal translation model comprises an image transformer, a second text encoder, and a second text decoder, and inputting the source text and the source associated image into the pre-generated target multimodal translation model to obtain the target translation text output by the target multimodal translation model comprises: inputting the source associated image into the image transformer to obtain a target prompt text output by the image transformer; inputting the source text and the target prompt text into the second text encoder to obtain a second source text feature output by the text encoder; and inputting the second source text feature into the second text decoder to obtain the target translation text output by the text decoder.
However, Parida discloses a multimodal neural machine translation system and method (Abstract), comprising the following: a target multimodal translation model comprises an image transformer (R-CNN with Resnet-101-C4), a second text encoder (mBART encoder), and a second text decoder (mBART decoder) (3 Description of the MMT System, pg. 32 and 33), and inputting source text and source associated image into a pre-generated target multimodal translation model to obtain a target translation text output by the target multimodal translation model comprises: inputting the source associated image into the image transformer to obtain a target prompt text output (object tags) by the image transformer (3. Description of the MMT system, pg. 32 and 33); inputting the source text and the target prompt text into a second text encoder (mBART encoder) to obtain a second source text feature output by the text encoder (3. Description of the MMT system, pg. 32 and 33); and inputting the second source text feature into the second text decoder (mBART decoder) to obtain the target translation text output by the text decoder (3. Description of the MMT system, pg. 32 and 33).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kwon, Kwon, Gronroos and Meng in the same way that Parida’s invention has been improved to achieve the following predictable results for the purpose of performing enhanced multimodal translation for English to Bengali: wherein the target multimodal translation model further comprises the image transformer, a second text encoder, and a second text decoder, and inputting the source text and the source associated image into the pre-generated target multimodal translation model to obtain the target translation text output by the target multimodal translation model comprises: inputting the source associated image into the image transformer to obtain a target prompt text output by the image transformer; inputting the source text and the target prompt text into the second text encoder to obtain a second source text feature output by the text encoder; and inputting the second source text feature into the second text decoder to obtain the target translation text output by the text decoder.
Claim(s) 6, 16 and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over in view of Kwon et al. (“A text-based visual context modulation neural model for multimodal machine translation”) (“Kwon”) in view of Gronroos et al. (The MeMAD Submission to the WMT18 Multimodal Translation Task) (“Gronroos”), and further in view of Meng et al. (US 2022/0245365) (“Meng”), and further in view of Parida et al. (“Multimodal Neural Machine Translation System for English to Bengali”), and further in view of Hu et al. (US 2023/0120631) (“Hu”).
For claims 6, 16 and 22, the combination of Kwon, Gronroos, Meng and Parida further discloses wherein the undetermined multimodal translation model comprises the image encoder, the second text encoder, and the first second decoder (Parida, 3. Description of the MMT system and 4.3 Translation using image and text modality, pg. 32 – 34), and the target multimodal translation model is generated by: obtaining the sample data (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609) (Parida, 1. Introduction, 4.1 BVG Dataset and 4.2 Translation using image and text modality, pg. 31, 32 and 34); and cyclically performing a first training step according to the sample data (Gronroos, 4.3 Training, pg. 608 and 609) (Parida, 4.1 BVG Dataset and 4.2 Translation using image and text modality, pg. 34), and using the trained undetermined multimodal translation model as the target multimodal translation model (Meng, Fig.1; [0040] [0150] [0186]) (Parida, Figure 2), wherein the second training step comprises: determining an undetermined prompt text (object tags), a second undetermined source text, and a second undetermined target text according to the sample data (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609) (Parida, 1. Introduction, 4.1 BVG Dataset and 4.2 Translation using image and text modality, pg. 31, 32 and 34); inputting the second undetermined source text and the undetermined prompt text into the second text encoder to obtain a second text sample feature output by the second text encoder (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609) (Parida, 1. Introduction, 3 Description of the MMT System, 4.1 BVG Dataset and 4.2 Translation using image and text modality, pg. 31- 34); and inputting the second text sample feature into the second text decoder to obtain a second translation text output by the second text decoder (Gronroos, Table 1, 1. Introduction, 4 Experiment 3: Multi-modal Transformer, 4.3, Training, pg. 603, 606, 608 and 609) (Parida, 1. Introduction, 3 Description of the MMT System, 4.1 BVG Dataset and 4.2 Translation using image and text modality, pg. 31- 34).
Yet, the combination of Kwon, Gronroos, Meng and Parida fails to teach the following: cyclically performing the first training step until it is determined that the trained undetermined multimodal translation model meets a first preset stop condition; determining a second loss value according to the second translation text and the second undetermined target text, and when it is determined, according to the first loss value, that the undetermined multimodal translation model does not meet the first preset stop iteration condition, updating a parameter of the undetermined multimodal translation model according to the first loss value to obtain a trained undetermined multimodal translation model, and using the trained undetermined multimodal translation model as a new undetermined multimodal translation model.
However, Hu discloses a neural network training method (Abstract), comprising the following: determining a first loss value according to a result label value of training data and a result output by a neural network model ([0172]), and when it is determined, according to the first loss value, that the undetermined neural network model does not meet the first preset stop iteration condition, updating a parameter of the undetermined neural network model according to the first loss value to obtain a trained undetermined neural network model ([0172] [0193] [0194] [0243 – 0247]) and using the trained undetermined multimodal translation model as a new undetermined multimodal translation model ([0239 – 0242] [0248]). Furthermore, Hu discloses cyclically performing the first training step until it is determined that the trained undetermined multimodal translation model meets a first preset stop condition ([0243 - 0248]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kwon, Gronroos, Meng and Parida in the same way that Hu’s invention has been improved to achieve the following, predictable results for the purpose of efficiently training the undetermined multimodal translation model to achieve a desired translation quality: cyclically performing the first training step until it is determined that the trained undetermined multimodal translation model meets a first preset stop condition; determining a first loss value according to an output, e.g. second translation text, of the undetermined multimodal translation model and a result label value, e.g. second undetermined target text, of training data, and when it is determined, according to the first loss value, that the undetermined multimodal translation model does not meet a first preset stop iteration condition, updating a parameter of the undetermined multimodal translation model according to the first loss value to obtain a trained undetermined multimodal translation model, and using the trained undetermined multimodal translation model as a new undetermined multimodal translation model.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SONIA L GAY whose telephone number is (571)270-1951. The examiner can normally be reached Monday-Friday 9-5 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SONIA L GAY/Primary Examiner, Art Unit 2657