DETAILED ACTION
This examination is in response to the communication filed on 01/08/2025. Claims 1-20 are currently pending, where claims 1, 19 and 20 are independent.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/082025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities: page 18, lines 2-9 makes references to “images” which are not part of the specification nor clearly mapped to the drawings.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Independent claims 1, 19 and 20 recite “an input device configured for receiving generating a user input data…” making it unclear whether then input device receives and/or generates the user input data. For purposed of Examination, the limitation is interpreted as being “an input device configured for receiving or generating a user input device”.
Claims 2-18 variously depend from independent claim 1 and therefore are rejected for the same reasons a claim 1.
Claim 11 recites the limitation "the plurality of AI modules" in lines 7-8. There is insufficient antecedent basis for this limitation in the claim.
Claim 11 recites the limitation "the plurality of user devices” in line 13. There is insufficient antecedent basis for this limitation in the claim.
Claim 18 recites the limitation "the modified translation data” in line 12. There is insufficient antecedent basis for this limitation in the claim
Double Patenting
Applicant is advised that should claim 19 or 20 be found allowable, claims 2 and 15 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 3, 5, 9, 10, 12-15 and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Pham et al. (US 2022/0391603 A1; herein “Pham”).
Regarding claim 1, Pham teaches a real-time language translation system (Figs. 6L-6O illustrates a side-by-side conversation wherein real-time language translation is performed between English and Spanish audio inputs; Under a broadest reasonable interpretation “real-time” is interpreted as including during a conversation, as opposed to after completion of the conversation) embodied in a physical device (Fig. 6M, tablet 600), wherein the real-time language translation system comprises:
an input device configured for receiving a user input data representing a linguistic input from a user, wherein the linguistic input corresponds to a user language (Fig. 3 I/O interface 330; Microphone 113; Fig. 6A microphone 606DFig. 6K, audio input “Hi, How are you?” and ¶[0215] teaches “At FIG. 6K, electronic device 600 receives an audio input of a user saying ‘Hi, how are you?’”);
a processing device configured for generating a translation data based on the user input data, wherein the translation data represents a translation of the linguistic input, wherein the generating is based on an AI module, wherein the processing device is communicatively coupled to the input device (FIG. 3 CPU(s) 310; FIG. 6A, tablet 600; and FIG. 6K ¶[0216] teaches “…electronic device 600 determines that the audio input is in the first language (e.g., English) and displays transcription 636A of the audio input in the first language and translation 636B of the audio input in the second language on the left side of side-by-side conversation user interface 616”); and
a presentation device configured for presenting the translation data, wherein the presentation device is communicatively coupled to the processing device (FIG. 3, Display 340; Fig. 6A, touch-sensitive display 602 and FIG. 6K ¶[0216] teaches “…electronic device 600 determines that the audio input is in the first language (e.g., English) and displays transcription 636A of the audio input in the first language and translation 636B of the audio input in the second language on the left side of side-by-side conversation user interface 616”).
Regarding claim 3, Pham teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Pham further teaches the user input data comprises at least one of an audio data corresponding to an audio input from the user and a textual data corresponding to a text input from the user (¶[0203] teaches “Electronic device 600 depicts translation user interface 604…for inputting content (e.g., text and/or handwritten letters) to be translated…Translation user interface 604 also includes audio input option 606D that is selectable to provide an audio input (e.g., a spoken input) for translation”).
Regarding claim 5, Pham teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Pham further teaches the translation data comprises a GUI data configured for presenting a GUI on the presentation device (Figs 6A-6F, translation user interface 604), wherein the GUI data comprises an activation data representing an activation parameter configured to initiate generating the user input data by the input device associated with the user (FIG. 6F, elements 618D and 618E and ¶[0209] teaches “…Side-by-side conversation user interface 616 also includes first audio option 618D that is selectable to cause electronic device 600 to be configured to receive an audio input in the first language…”).
Regarding claim 9, Pham teaches all of the elements of claim 3 (see detailed element mapping above). In addition, Pham further teaches the audio data comprises an audio representation data corresponding to a representation of the audio data (¶[0203] teaches “Electronic device 600 depicts translation user interface 604…for inputting content (e.g., text and/or handwritten letters) to be translated…Translation user interface 604 also includes audio input option 606D that is selectable to provide an audio input (e.g., a spoken input) for translation” ),
wherein the audio representation data comprises at least one of a raw audio data (¶[0059] teaches “Audio circuitry 110…provide an audio interface between a user and device 100…Audio circuitry 110 also receives electrical signals converted by microphone 113 from sound wave” the electrical signal is interpreted as raw audio data), a Fourier transformation data, a spectrogram data, a mel-frequency cepstal coefficient data and a vector embedding data, wherein the raw audio data corresponds to an unprocessed audio input (¶[0059] teaches “Audio circuitry 110…provide an audio interface between a user and device 100…Audio circuitry 110 also receives electrical signals converted by microphone 113 from sound wave” the electrical signal are unprocessed audio input), wherein the Fourier transformation data of the audio data corresponds to converting an audio waveform associated with the audio data from a time domain to a frequency domain, wherein the spectrogram data corresponds to a visual representation of a spectrum of frequencies associated with the audio data, wherein the mel-frequency cepstal coefficient data represents a short-term power spectrum of the audio input, wherein the vector embedding data corresponds to a vector representation of the audio data (the at least one of language makes the other limitations optional).
Regarding claim 10, Pham teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Pham further teaches the AI module comprises a plurality of AI modules (¶[0290] teaches “…a determination can be made (e.g., based on one or more machine learning models), that translation 826B of transcription 826A does not satisfy a threshold confidence score or level…”).
Regarding claim 12, Pham teaches all of the elements of claim 10 (see detailed element mapping above). In addition, Pham further teaches the user input data comprises a plurality of user input data corresponding to a plurality of users (As shown in the Face-to-face conversation illustrated in Fig. 6Z, the conversational input corresponds to two users),
wherein the user language comprises a plurality of user languages corresponding to the plurality of user input data (Fig. 6Z, Spanish and English, wherein each of the plurality of user languages comprises a plurality of user dialects (the Spanish and English languages inherently comprise a plurality of user dialects),
wherein the plurality of AI modules further comprises at least one of a user ID AI model (the “at least one of” language renders this element optional), a language ID AI model (¶[0242] teaches “in accordance with a determination that the automatic language detection setting (e.g., 628C) is enabled, …a general listening mode of the computer system in which the computer system is configured to receive an audio input in a plurality of languages (e.g., is configured to transcribe audio inputs in the plurality of languages”) and a dialect ID AI model (the “at least one of” language renders this element optional), wherein the user ID AI model is configured for generating a user ID data corresponding to each of the plurality of users based on the user input data (the “at least one of” language renders this element optional), wherein the language ID AI model is configured for generating a user language ID data corresponding to each of the plurality of user languages based on the user input data (¶[0242] teaches “in accordance with a determination that the automatic language detection setting (e.g., 628C) is enabled, …a general listening mode of the computer system in which the computer system is configured to receive an audio input in a plurality of languages (e.g., is configured to transcribe audio inputs in the plurality of languages” and ¶[0213] teaches “Option 628Ccorresponds to an automatic language detection setting, and is selectable to enable and/or disable the automatic language detection setting. When automatic language detection is enabled, electronic device 600 receives an audio input, and automatically determines whether the input is in the first language or in the second language. ), wherein the dialect ID AI model is configured for generating a user dialect ID data corresponding to each of the plurality of user dialects based on the user input data (the “at least one of” language renders this element optional).
Regarding claim 13, Pham teaches all of the elements of claim 12 (see detailed element mapping above). In addition, Pham further teaches the translation data comprises a GUI data configured for presenting a GUI on the presentation device, wherein the GUI data comprises a language detection button data corresponding to initiation of generating the user language ID data based on the language ID AI model (FIG. 6I element 628C and ¶[0213] teaches “Option 628C corresponds to an automatic language detection setting, and is selectable to enable and/or display the automatic language detection setting” ).
Regarding claim 14, Pham teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Pham further teaches the user input data comprises an audio data corresponding to an audio input from the user (¶[0203] teaches “Translation user interface 604 also includes audio input option 606D that is selectable to provide an audio input (e.g., a spoken input) for translation.” ),
wherein the audio data comprises an audio characteristic data corresponding to a characteristic associated with the audio input, wherein the audio characteristic data comprises an audio energy level data corresponding to an energy level associated with the audio input, wherein the energy level further corresponds to a numerical value associated with the audio data (As noted above, Pham teaches acquiring a spoken input, a spoken input inherently includes audio characteristic data corresponding to an energy level. In addition, Pham specifically teaches the continuous listening mode includes detecting volume level (which is numerical value associated with the audio inputs energy level) to detect when the speech input has completed (See ¶[0213]). Accordingly, Pham teaches that the spoken input includes audio characteristic data comprising energy level data corresponding to a numerical value).
Regarding claim 15, Pham teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Pham further teaches a storage device configured for storing each of the user input data and the translation data associated with the user input data (¶[0056] teaches “Memory 102 optionally includes…such as one or more magnetic disk storage devices” ).
Regarding claim 20, Pham teaches a real-time language translation system (Figs. 6L-6O illustrates a side-by-side conversation wherein real-time language translation is performed between English and Spanish audio inputs; Under a broadest reasonable interpretation “real-time” is interpreted as including during a conversation, as opposed to after completion of the conversation) embodied in a physical device (Fig. 6M, tablet 600), wherein the real-time language translation system comprises:
an input device configured for receiving generating a user input data representing a linguistic input from a user, wherein the linguistic input corresponds to a user language (Fig. 3 I/O interface 330; Microphone 113; Fig. 6A microphone 606DFig. 6K, audio input “Hi, How are you?” and ¶[0215] teaches “At FIG. 6K, electronic device 600 receives an audio input of a user saying ‘Hi, how are you?’”);
a processing device configured for generating a translation data based on the user input data, wherein the translation data represents a translation of the linguistic input, wherein the generating is based on an AI module, wherein the processing device is communicatively coupled to the input device (FIG. 3 CPU(s) 310; FIG. 6A, tablet 600; and FIG. 6K ¶[0216] teaches “…electronic device 600 determines that the audio input is in the first language (e.g., English) and displays transcription 636A of the audio input in the first language and translation 636B of the audio input in the second language on the left side of side-by-side conversation user interface 616”); and
a presentation device configured for presenting the translation data, wherein the presentation device is communicatively coupled to the processing device (FIG. 3, Display 340; Fig. 6A, touch-sensitive display 602 and FIG. 6K ¶[0216] teaches “…electronic device 600 determines that the audio input is in the first language (e.g., English) and displays transcription 636A of the audio input in the first language and translation 636B of the audio input in the second language on the left side of side-by-side conversation user interface 616”); and
a storage device configured for storing each of the user input data and the translation data associated with the user input data (¶[0056] teaches “Memory 102 optionally includes…such as one or more magnetic disk storage devices”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Pham as applied to claim 1 above, and further in view of Voss et al. (US 2022/0199071 A1; herein “Voss”).
Regarding claim 11, Pham teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Pham further teaches the user input data comprises an audio data corresponding to an audio input from the user, wherein the audio data is in user language, wherein the audio data comprises an audio characteristic data corresponding to a characteristic associated with the audio input (Fig. 6M and ¶P[0217]-[0218] teaches “…because the continuous listening setting is enabled, electronic device 600 remains in the listening mode…and received audio input of a user saying ‘I’m a little hungry’...based on the automatic detection setting being enabled, electronic device 600 determines that the audio input is in the first language…”),
wherein the audio data further comprises at least one of a speech data (Fig. 6M input audio “I’m a little hungry” and ¶[0217] teaches “…because the continuous listening setting is enabled, electronic device 600 remains in the listening mode…and received audio input of a user saying ‘I’m a little hungry’.”) and a noise data (the “at least one” language renders this element optional), wherein the speech data represents a word spoken by the user (Fig. 6M input audio “I’m a little hungry” and ¶[0217] teaches “…because the continuous listening setting is enabled, electronic device 600 remains in the listening mode…and received audio input of a user saying ‘I’m a little hungry’.”) and a noise data (the at least one language renders this element optional), wherein the noise data represents the audio data excluding the speech data (the “at least one” language renders this element optional),
wherein the plurality of AI modules comprises a voice activity detection AI module configured for detecting the speech data from the audio data, wherein the detection is based on the audio characteristic data associated with the audio input (¶[0213] teaches “When the continuous listening setting is enabled…will cause electronic device to continuously listen for audio inputs…when no audio input and/or audio input surpassing a threshold volume is not detected for a threshold period of time), electronic device 600 will automatically stop listening for audio input…”)
wherein the plurality of AI modules comprises a neural machine translation AI module (¶[0290] teaches “…a determination can be made (e.g., based on one or more machine learning models), that translation 826B of transcription 826A…”)further configured for translating the tokenized text data representing the first user language associated with the first user into a translated tokenized text data representing the second user language associated with the second user (Fig. 6N illustrates that the audio input “I’m a little hungry” shown in Fig. 6M, is tokenized into English text 640A and translated into tokenize text in the second language, i.e., the translated text data 640B; ), wherein the neural machine translation AI module is further configured for translating translated tokenized text data representing the second user language to the tokenized text data representing the first user language (Fig. 6N illustrates that the audio input “Bien, Y Tu?” shown in Fig. 6L, is tokenized into Spanish (e.g., the second language) text 638A and translated into tokenize text in the first language, i.e., the translated text data 638B),
wherein the plurality of AI modules further comprises a text to speech AI module configured for generating a translated audio data corresponding to a translated user language from the audio data (¶[0213] teaches “…When autopay is enabled…a translation of the input (e.g., in the other of the second language and/or first language) is automatically spoken aloud by the electronic device 600”),
wherein a first user device associated with a first user and a second user device associated with a second user (Under a broadest reasonable interpretation, the first and second user device need not be different devices as long as portions of the device interface are associated with different users; The face-to-face conversation shown in Fig. 6R teaches that different portions of the device interface are associated with a first and second user, i.e., first and second language. Specifically, ¶[0222] teaches “Accordingly, in face-to-face conversation user interface 652, components that correspond to the first language ( e.g., first language selector 618A, first audio input option 618D) are displayed in a first orientation, and components that correspond to the second language ( e.g., second language selector 618B, second audio input option 618E) are displayed in a second orientation (e.g., a second orientation upside down relative to the first orientation” and ¶[0223] teaches “However, face-to-face conversation user interface 652 includes two audio input options 618D and 618E, even when automatic detection is enabled (e.g., so that one user does not have to reach across a table and/or across device 600 to select the audio input option).” Thus, there are specific controls and portions of the GUI which are associated with the first user and specific controls and portions of the GUI which are associated with the second user)
wherein the presentation device comprises a display device configured to display each of the text data corresponding to the tokenized text data and the translated text data corresponding to the translated text data, wherein the presentation device comprises a speaker configured to present the translated audio data (Fig. 6N illustrates that the audio input “I’m a little hungry” shown in Fig. 6M, is displayed a text in the English language 640A, i.e., the tokenized text data, and the translated text data 640B; and ¶[0213] teaches “…When autopay is enabled…a translation of the input (e.g., in the other of the second language and/or first language) is automatically spoken aloud by the electronic device 600”).
Although Pham teaches converting audio input into text in a first language and translating the text into a second language, a GUI for displaying the translated conversation, and the use of one or more machine learning algorithms/models, Pham fails to disclose any specificity as what machine learning models are utilizes. Therefore, Pham fails to specifically disclose (1) wherein the above noted functional it performed using machine learning algorithms, (2) wherein the plurality of AI modules further comprises an automatic speech recognition AI module configured for generating a text data based on the speech data, wherein the generating is based on Connectionist Temporal Classification Beam Search algorithm, wherein the automatic speech recognition AI module is further configured for tokenizing the text data to a tokenized text data associated with the text data, (3) wherein the plurality of user devices comprises a first user device associated with a first user and a second user device associated with a second user.
Voss teaches an automatic speech recognition AI module configured for generating a text data based on the speech data, wherein the generating is based on Connectionist Temporal Classification Beam Search algorithm, wherein the automatic speech recognition AI module is further configured for tokenizing the text data to a tokenized text data associated with the text data (¶[0078] teaches “Processes in accordance with many embodiments of the invention can utilize beam search over grapheme or phoneme probability vectors such as CTC tokens to generate predictions…” ).
Pham differs from the claimed invention, as defined in claim 11, in that Pham fails to explicitly disclose that the ASR functionality is achieved using an AI module based on a CTC token beam search algorithm. Speech recognition using an AI module based on a CTC token beam search algorithm is known in the art as evidenced by Voss. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention to have implemented the ASR functionality in system taught by Pham with a CTC token beam search algorithm as taught by Voss as it merely constitutes the combination of known processes/methods to achieve the predictable result of transcribing an audio signal to text in a given language (Voss,¶[0078]).
Claims 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Pham as respectively applied to claims 13 and 5 above, and further in view of Park et al. (US 2024/0104311 A1; herein “Park”).
Regarding claim 16, Pham teaches all of the elements of claim 15 (see detailed element mapping above). In addition, Pham further teaches the storage device further configured for retrieving the user input data and the associated translation data (¶[0236] teaches “…the first region includes a plurality transcription/translation pairs corresponding to previous inputs received in the first language…the second region includes a plurality of transcription/translation pairs corresponding to previous inputs received in the second language” displaying previous received inputs in interpreted as retrieving from memory/storage the previous transcription/translation pairs”).
However, Pham fails to disclose a communication device configured for transmitting each of the user input data and the translation data associated with the user input data to an external server.
Park teaches a communication device configured for transmitting each of the user input data and the translation data associated with the user input data to an external server (¶[0029] teaches “A user’s speech is captured and transmitted to a remote server over a network where it is analyzed and translated before being sent back to a client device.”; ¶[0064] teaches “The speech input 602 may also be transmitted to a remote computing device (e.g., cloud server) 616 to compute
a translation of the speech input 602.”, and ¶[0073] teaches “the correction module 628 may select the on-device translation (generated via NLU 620) or the cloud-based translation 624 to be output via the mobile device. In some aspects, the correction module 628 may supply the on-device translation for output as well as a correction determined based on the cloud-based translation 624.”).
Pham differs from the claimed invention, as defined in claim 16, in that Pham fails to disclose transmitting the user input data and the translation data associated with the user input to an external server. Transmitting user input and translation data from a mobile device to an external server for accuracy comparison/correction is known in the art as evidenced by Park. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the translation system taught by to include cloud-based translation comparison as taught by Park as it merely constitutes that combination of known processes to achieve the predictable result of improving the accuracy of the on device translation model.
Regarding claim 17, Pham teaches all of the elements of claim 13 (see detailed element mapping above). In addition, Pham further teaches the translation data is presented on a user display device associated with the user device (See translation user interface 604). However, Pham fails to explicitly disclose wherein the user device comprises a user input device configured for receiving a feedback data corresponding to a feedback of the translation data, wherein the user device further comprises a user-processing device configured for generating a modified translation data based on the feedback data, wherein the user display device is further configured to present the modified translation data, wherein the generating is based on the AI model, wherein the communication device is further configured for receiving each of the feedback data and the modified translation data.
Park teaches a mobile translation device that includes, inter alia, a user input device configured for receiving a feedback data corresponding to a feedback of the translation data (¶[0091] teaches “At block 886, a determinization of whether negative user feedback is detected” and ¶[0094] teaches “…negative feedback from a microphone (e.g., an ear bud) may be determined based on keywords…a motion sensor may detect a tilted head…the negative feedback may be recognized with camera based gesture recognition”), and
a user-processing device configured for generating a modified translation data based on the feedback data, wherein the user display device is further configured to present the modified translation data, wherein the generating is based on the AI model, wherein the communication device is further configured for receiving each of the feedback data and the modified translation data (¶[0091] teaches “If negative feedback has been detected, at block 888, a corrected translation may be determined based on the cloud-based translation 812” ).
Pham differs from the claimed invention, as defined in claim 17, in that Pham fails to disclose that the receiving user feedback regarding the translation and correcting the translation based on the feedback. Using feedback regarding translation accuracy and correcting the translations based on the feedback is known in the art as evidenced by Park. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the translation system taught by to include user feedback as taught by Park as it merely constitutes that combination of known processes to achieve the predictable result of improving the accuracy of the translation model.
Regarding claim 18, Pham teaches all of the elements of claim 5 (see detailed element mapping above). In addition, Pham further teaches the user comprises a plurality of users (The face-to-face conversations mode comprises a plurality of users, i.e., English speaker and Spanish speaker),
wherein the GUI data comprises a user screen data corresponding to a screen presented on a user display device associated with the user device, wherein the user screen data comprises a multi button data corresponding to a plurality of buttons, wherein each of the plurality of buttons is associated with each of the plurality of users (See Fig.6R, includes multiple buttons, i.e., microphone 618D and 618E each associated with a different user),
wherein the user display device is further associated with the user device comprising a user input device configured for receiving a feedback data corresponding to a feedback corresponding to the multi button data (Fig. 6T, user input 658 and ¶¶[0224]-[0225] teaches that in response to user input corresponding to selection of one of the multiple microphone buttons, the other button is greyed out or otherwise indicated as disabled.),
wherein the user device further comprises a user-processing device configured for generating a modified multi button data based on the feedback data (Fig. 6U, disabled microphone 618E and ¶¶[0224]-[0225] teaches that in response to user input corresponding to selection of one of the multiple microphone buttons, the other button is greyed out or otherwise indicated as disabled.).
However, Pham fails to disclose the user display device is further configured to present the modified translation data.
Park teaches wherein the user display device is further configured to present the modified translation data (¶[0091] teaches “If negative feedback has been detected, at block 888, a corrected translation may be determined based on the cloud-based translation 812”).
Pham differs from the claimed invention, as defined in claim 18, in that Pham fails to disclose that the receiving user feedback regarding the translation and correcting the translation based on the feedback. Using feedback regarding translation accuracy and correcting the translations based on the feedback is known in the art as evidenced by Park. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the translation system taught by to include user feedback as taught by Park as it merely constitutes that combination of known processes to achieve the predictable result of improving the accuracy of the translation model.
Allowable Subject Matter
Claims 2, 4, 6-8 and 19 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) set forth in this Office action.
Regarding claim 19, Pham teaches a real-time language translation system (Figs. 6L-6O illustrates a side-by-side conversation wherein real-time language translation is performed between English and Spanish audio inputs; Under a broadest reasonable interpretation “real-time” is interpreted as including during a conversation, as opposed to after completion of the conversation) embodied in a physical device (Fig. 6M, tablet 600), wherein the real-time language translation system comprises:
an input device configured for receiving generating a user input data representing a linguistic input from a user, wherein the linguistic input corresponds to a user language (Fig. 3 I/O interface 330; Microphone 113; Fig. 6A microphone 606DFig. 6K, audio input “Hi, How are you?” and ¶[0215] teaches “At FIG. 6K, electronic device 600 receives an audio input of a user saying ‘Hi, how are you?’”);
a processing device configured for generating a translation data based on the user input data, wherein the translation data represents a translation of the linguistic input, wherein the generating is based on an AI module, wherein the processing device is communicatively coupled to the input device (FIG. 3 CPU(s) 310; FIG. 6A, tablet 600; and FIG. 6K ¶[0216] teaches “…electronic device 600 determines that the audio input is in the first language (e.g., English) and displays transcription 636A of the audio input in the first language and translation 636B of the audio input in the second language on the left side of side-by-side conversation user interface 616”); and
a presentation device configured for presenting the translation data, wherein the presentation device is communicatively coupled to the processing device (FIG. 3, Display 340; Fig. 6A, touch-sensitive display 602 and FIG. 6K ¶[0216] teaches “…electronic device 600 determines that the audio input is in the first language (e.g., English) and displays transcription 636A of the audio input in the first language and translation 636B of the audio input in the second language on the left side of side-by-side conversation user interface 616”).
However, Pham fails to disclose or suggest that the physical device is configured to be affixed on a user device associated with the user, and wherein the physical device comprises at least one of a mobile case and an enclosed backpack.
Claim 2 is essentially a duplicate of independent claim 19, and claims 2, 4, 6-8 depend from claim 2. Therefore, claims 2, 4 and 6-8 contain the allowable subject matter of claim 19 noted above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PENNY L CAUDLE whose telephone number is (703)756-1432. The examiner can normally be reached M-Th 8:00 am to 5:00 pm eastern.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PENNY L CAUDLE/Examiner, Art Unit 2657
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657