DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application is being examined under the pre-AIA first to invent provisions.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-14 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-2, 4-8, 10, and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blair, US 2003/0050776 A1 (previously cited), in view of Lemelson et al., US 6,028,514 A (previously cited and hereinafter Lemelson, also cited in IDS received 5/4/2024), and further in view of Cohen, US 2010/0121629 A1.
Regarding claim 1, Blair discloses a message capturing device that records spoken messages and automatically translates the recorded messages into text format (see abstract and figure 1). Blair teaches various embodiments of the message capturing system, for example a combination handheld recording and translation device (e.g., a cell phone) records spoken messages and automatically translates the recorded messages into text format (see Blair, ¶ 0011-0012, 0021, and 0030, figure 1, items 20, 25, and 30, and figure 3, item 25).
Herein, Blair anticipates a “communication system” by teaching a message capturing system including a combination handheld recording and translation device, such as a cell phone, and a remote presentation device, such as a desktop computer, handheld device, cell phone, etc. (see Blair, ¶ 0011-0012, 0021, and 0030, figure 1, items 20, 25, 30, and 40, and figure 3, items 25 and 40B-C),
“comprising:
a loudspeaker” by teaching the combination handheld recording and translation device, such as a cell phone, and the cell phone is well-known to comprise a loudspeaker at least for telephonic communication (see Blair, ¶ 0011-0012 and 0030, figure 1, items 20, 25, and 30, and figure 3, item 25); and
“a microphone configured to measure the ambient environment and generate a microphone signal” by teaching that the microphone of the cell phone (i.e., Blair’s recording device) is used to detect voice commands that direct the cell phone to start recording a voice message, and further teaches that the cell phone stops recording when the user’s voice is no longer detected or another voice command is recognized, wherein the features to detect voice commands to start and stop recording anticipates a microphone generating a microphone signal when only the ambient environment is audible (see Blair, ¶ 0012, 0015-0016, and 0027 and figure 1, items 20 and 25, and figure 3, items 24-25).
“a memory that is configured to store program code as instructions” by teaching the cell phone with voice recognition software (see Blair, ¶ 0011-0012 and 0015-0017); and
“a processing unit that executes the instructions to perform operations” by teaching that the voice recognition software is executed on a processor (see Blair, ¶ 0011-0012, 0015, and 0017).
However, Blair does not appear to teach the operations comprising “receiving the” [accelerometer] “sensor data” and/or “analyzing the sensor data to generate a sensor value”.
Lemelson teaches a warning system with a portable, personal warning unit for a user to carry (see Lemelson, abstract, column 3, lines 4-5, column 11, lines 1-5, and figures 1-2, unit 12,), with a motion detector (e.g., an accelerometer) that is used to detect ‘abnormal rapid changes in movement’ of the user carrying the personal warning unit (see Lemelson, column 11, lines 24-30 and lines 49-52, and figure 2, units 42, 44, and 53). Lemelson also teaches speech recognition features to detect specific preprogrammed speech (see Lemelson, column 4, lines 60-67), and also teaches sound recognition features to detect certain types of sounds that indicate an emergency, such as riot sounds, gunshots, or other loud noises (see Lemelson, column 4, line 67 – column 5, line 4). Specifically, Lemelson teaches the combination of motion analysis and sound analysis to determine alarm conditions (see Lemelson, column 15, lines 19-22). It would have been obvious to one of ordinary skill in the art at the time of the invention to modify Blair with the teachings of Lemelson for the purpose of providing automated emergency features for cell phone users (see Blair, ¶ 0011-0012, 0021, and 0030, and figure 3, unit 25, in view of Lemelson, column 9, line 63 – column 10, line 21, column 11, lines 24-59, and figures 1-2, unit 12).
Therefore, the combination of Blair and Lemelson makes obvious the communication system comprising:
“receiving the sensor data;” because the processor of the device receives accelerometer sensor data and/or signals (see Blair, ¶ 0030 in view of Lemelson, column 11, lines 49-52, column 15, lines 14-15, figure 2, unit 53, and figure 4A, step 163);
“analyzing the sensor data to generate a sensor value” because Lemelson makes it obvious to analyze the sensor data in order to determine if the sensor output indicates a rapid and/or unusual change, and Lemelson makes obvious the use of fuzzy logic to determine sensor values (see Lemelson, column 5, lines 16-38, column 11, lines 49-56, and column 15, lines 14-42);
“receiving a first signal, wherein the first signal is the microphone signal or a modified version of the microphone signal” because Blair teaches that the cell phone is listening for voice commands in order to perform certain tasks, such as starting and stopping a voice message recording, and the cell phone records the voice message, wherein the received first signal is anticipated by any one of the received voice command and/or voice message transduced from the cell phone’s microphone (see Blair, ¶ 0011-0012 and 0015-0017);
“detecting speech in the first signal using the signature sound ID system” because Blair teaches that speech is detected in the received microphone signal in order to starting and stopping a voice message recording (see Blair, ¶ 0015-0016);
“converting the speech to text” because Blair teaches that the cell phone translates, or converts, the speech to text (see Blair, ¶ 0017 and 0027, figure 1, items 20, 22, 25, 30, and 32, and figure 3, items 25 and 32); [and]
“inserting the text into a text transcript as the speech is converted to text” because Blair teaches known voice-recognition software, that allows user control over punctuation and message formatting where this reads on inserting text into a text transcript (i.e., a text file) while responding to both user voice commands for punctuation and/or message formatting embedded in the voice message (see Blair, ¶ 0015-0017 and 0020).
However, the combination of Blair and Lemelson do not appear to teach the features for “translating the text transcript to a prechosen language” and “converting the translated text transcript to a translated speech signal”.
Cohen teaches a translation platform for translating voice and/or text from a first language to a second language (see Cohen, abstract). Herein, Cohen teaches speech-to-text engines for converting human speech into text (see Cohen, ¶ 0026, 0033, and figure 1, unit 111), converting the text from one language into another (see Cohen, ¶ 0033 and figure 1, unit 112), and teaches text-to-speech engines to convert the translated text into an audio signal (see Cohen, ¶ 0031 and figure 1, unit 113). It would have been obvious to one of ordinary skill in the art at the time of the invention to modify the combination of Blair and Lemelson with the teachings of Cohen for the purpose of improving communication between users speaking different languages (see Blair, ¶ 0003-0005 in view of Cohen, ¶ 0005 and 0024-0025).
Therefore, the combination of Blair, Lemelson, and Cohen makes obvious the additional features for:
“generating a translated text transcript by translating the text transcript to a prechosen language” because Cohen makes obvious that a user selects a second language for the text translation (see Cohen, ¶ 0025-0026 and 0033, and figure 1, unit 112);
“converting the translated text transcript to a translated speech signal” because Cohen makes obvious that the translated text is synthesized into the second language so that another person can understand the content in their language (see Cohen, ¶ 0023-0026 and 0031, and figure 1, unit 113);
“sending the text transcript and the translated text transcript to memory” because Blair teaches that the user designates a destination of the text transcript, such as saving a soft copy, or word processing document, where the soft copy is stored on a memory in a computer (see Blair, ¶ 0003, 0014-0015, 0021, and 0026-0027, figure 1, items 32, 40, and 42, and figure 3, items 32, 38, 40B-C, and 42B-C), and Cohen further makes it obvious to store the translated text transcript (see Blair, ¶ 0003 in view of Cohen, ¶ 0025); and
“sending the translated speech signal to the speaker” because Cohen makes it obvious to output the synthesized voice signal in another language so that another person can understand the content in their language (see Cohen, ¶ 0023-0024 and 0031).
Regarding claim 2, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 1, wherein a remote device generates a translated text transcript and sends the translated text transcript to the communication system” because the user’s wireless communication device sends the message to a central server to translate the text and to communicate the translated text and voice to the receiving party (see Blair, ¶ 0011, 0014, 0026-0027, and 0030, in view of Cohen, ¶ 0023-0026 and 0029-0033).
Regarding claim 4, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 1, wherein prior to the operations, of detecting speech, converting the speech, inserting the text and sending the text, a speech to text mode is activated” (emphasis added, see claim objections above) by teaching user inputs to start recording a voice memo for translation to text (see Blair, ¶ 0014, 0017, and 0027 and figure 3, items 25-26, and also see Cohen, ¶ 0023-0026).
Regarding claim 5, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 4, where the speech to text mode is manually activated” by teaching user input to send a voice message to a translation software (see Blair, ¶ 0014, 0017, and 0027).
Regarding claim 6, see the preceding rejection with respect to claim 4 above. The combination makes obvious the “communication system according to claim 4, where the speech to text mode is verbally activated” by teaching a user voice command to send a voice message to a translation software (see Blair, ¶ 0015-0017 and 0027).
Regarding claim 7, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 1, wherein the operation of sending the text transcript and the translated text transcript to memory is initiated manually” by teaching user input to send the text to a remote device (see Blair, ¶ 0014, 0017, and 0027, and see Cohen, ¶ 0023 and 0025-0026).
Regarding claim 8, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 1, wherein the operation of sending the text transcript and the translated text transcript to memory is initiated verbally” by teaching a user voice command to send the text to a remote device (see Blair, ¶ 0016-0017, and 0027, and see Cohen, ¶ 0023 and 0025-0026) .
Regarding claim 10, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 1, wherein the text transcript is a text message to a party, where the party is not the user of the communication device” because the user’s wireless communication device sends the message to a control server that communicates the text transcript to another (see Cohen, ¶ 0023-0026 and 0032).
Regarding claim 14, see the preceding rejection with respect to claim 1 above. The combination makes obvious the “communication system according to claim 1, wherein the communication device is a mobile phone” by teaching a cell phone (see Blair, ¶ 0011-0012 and 0027, figure 1, items 20, 25, and 30, and figure 3, item 25).
Claim(s) 3 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Blair, Lemelson, and Cohen as applied to claim 1 above, and further in view of Vensko et al., US 4,624,008 A (previously cited and hereinafter Vensko).
Regarding claim 3, see the preceding rejection with respect to claim 1 above. The combination of Blair, Lemelson, and Cohen makes obvious the communication system according to claim 1, wherein speech is detected by analysis of the input signal using the voice recognition software is executed on a processor (see Blair, ¶ 0011-0012, 0015, and 0017). The combination does not appear to teach storing a portion of the first signal in a circular buffer.
Vensko teaches an apparatus for automatic speech recognition (see Vensko, abstract). In particular, Vensko teaches a circular sentence buffer that allows the recognition algorithm to trail the input speech to the system and provide word recognition (see Vensko, column 4, line 30 – column 5, line 43). It would have been obvious to one of ordinary skill in the art at the time of the invention to modify the combination of Blair, Lemelson, and Cohen with the teachings of Vensko for the purpose of performing word recognition in a continuous speech recognition method (see Vensko, column 1, lines 40-56).
Therefore, the combination of Blair, Lemelson, Cohen, and Vensko makes obvious the “communication system according to claim 1, wherein the operations further include:
storing a portion of the first signal in a circular buffer, wherein the operation of detecting speech in the first signal is accomplished by using the signature sound ID system to analyze the portion” because Vensko makes obvious the use of a circular buffer to store portions of the input speech signal to then perform speech recognition on the stored portions (see Blair, ¶ 0011-0012, 0015, and 0017, in view of Vensko, column 4, line 30 – column 5, line 43, figure 4, and figure 5, unit 522).
Regarding claim 11, see the preceding rejection with respect to claim 1 above. The combination of Blair, Lemelson, and Cohen makes obvious the communication system according to claim 1, where the microphone signal is processed to perform speech recognition (see Blair, ¶ 0011-0012 and 0017). However, the combination does not teach the feature “wherein the modified microphone signal is generated by applying a gain to the microphone signal”.
Vensko teaches an apparatus for automatic speech recognition (see Vensko, abstract). Herein, Vensko teaches a microphone preamplifier circuit and a pre-emphasis amplifier that provides gain to the microphone signal to provide more gain in higher frequencies (see Vensko, column 2, lines 63-68, column 3, lines 35-44, figure 1, units 104 and 106, and figure 2, units 112, 200, and 202).
It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify the combination of Blair, Lemelson, and Cohen with the teachings of Vensko to try conventional speech recognition methods to perform the speech recognition and expect similar or better voice recognition results (see Vensko, column 3, lines 35-44).
Therefore, the combination of Blair, Lemelson, Cohen, and Vensko makes obvious the “communication system according to claim 1, wherein the modified microphone signal is generated by applying a gain to the microphone signal” because Vensko makes obvious to use conventional methods, such as using a pre-emphasis amplifier that provides gain to the microphone signal to provide more gain in higher frequencies (see Vensko, column 3, lines 35-44 and figure 2, units 112, 200, and 202).
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Blair, Lemelson, and Cohen as applied to claim 1 above, and further in view of Odinak et al., US 2003/0177009 A1 (previously cited and hereinafter Odinak).
Regarding claim 9, see the preceding rejection with respect to claim 1 above. The combination of Blair, Lemelson, and Cohen makes obvious the communication system according to claim 1 for sending the text transcript to a user’s email account (see Blair, ¶ 0030). However, the combination does not expressly teach the features of “adding a time stamp” to the text transcript prior to sending.
Odinak teaches a system and method for providing a message-based communications infrastructure for automated call center operation (see Odinak, abstract and figure 1). Similar to Blair, Odinak teaches automatic machine translation of speech to text annotation (see Odinak, ¶ 0070 and 0072-0073, and figure 4, units 54 and 56-57). Additionally, Odinak teaches adding time stamps to the transcript (see Odinak, ¶ 0079 and figure 5, unit 73).
It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify the combination of Blair, Lemelson, and Cohen with the teachings of Odinak for the purpose of keeping records related to the translated text transcript, such as keeping record of the time and/or date of the text translation transcript (see Blair, ¶ 0014-0017 and 0030 in view of Odinak, ¶ 0072 and 0079, figure 4, units 54 and 56-57, and figure 5, unit 73).
Therefore, the combination of Blair, Lemelson, Cohen, and Odinak makes obvious the “communication system according to claim 1, wherein the operations further include:
adding a time stamp to the text transcript prior to the operation of sending the text transcript and the translated text transcript to memory” because Blair teaches recording voice messages and translating the voice messages into a text transcript before sending the text transcript, and Odinak makes it obvious to add a time stamp to the audio and/or text transcript in order to keep records of the translation (see Blair, ¶ 0003, 0012, 0014-0017, and 0030, and Cohen, ¶ 0025, in view of Odinak, ¶ 0072 and 0079, figure 4, units 54 and 56-57, and figure 5, unit 73).
Claim(s) 12-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Blair, Lemelson, Cohen, and Vensko as applied to claim 11 above, and further in view of Lee, US 2005/0099400 A1 (previously cited)
Regarding claim 12, see the preceding rejection with respect to claim 11 above. The combination of Blair, Lemelson, Cohen, and Vensko makes obvious the communication system according to claim 11 with physical and voice input features. However, the combination does not appear to teach a touch-sensitive screen.
Lee discloses an apparatus and method for providing virtual graffiti, where providing virtual graffiti includes a touch-screen display for user touch input (see Lee, abstract, ¶ 0003, 0007, and 0022, figure 3A, units 300, 302, and 304, and figure 4A, units 400, 402, and 404). Lee teaches that the graffiti data entry method is used in the PALM OS using a touchpad adjacent to a touch-screen display (see Lee, ¶ 0005-0006 and figure 1, units 100, 102, and 108), and Lee teaches that prior art provided a virtual graffiti area to allow more efficient use of the touchscreen (see Lee, ¶ 0007, figure 3A, units 300, 302, and 304, and figure 3B, units 300 and 302).
It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify the combination of Blair, Lemelson, Cohen, and Vensko with the teachings of Lee for the purpose of providing an additional method of user input and improving the efficiency of a touchscreen display (see Blair, ¶ 0002, 0012, and 0014 in view of Lee, ¶ 0007, 0022, and 0037 and figures 3A-4B).
Therefore the combination of Blair, Lemelson, Cohen, Vensko, and Lee makes obvious the “communication system according to claim 11, wherein the communication device further comprises: a touch-sensitive screen” by teaching a recording and translating device, such as a cell phone or other ‘palm pilot’ device (see Blair, ¶ 0002, 0012, and 0027) and Lee makes obvious to use a touch-screen input of a cell phone to provide an additional method of user input (see Blair, ¶ 0014-0016 in view of Lee, ¶ 0007, 0022, and 0037 and figures 3A-4B).
Regarding claim 13, see the preceding rejection with respect to claim 12 above. The combination makes obvious the “communication system according to claim 12, wherein the communication device further comprises: a button” (see Blair, ¶ 0011 and 0027, and figure 3, items 25-26).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Holt et al., US 5,960,447 B1 (previously cited and hereinafter Holt), teaches a word tagging and editing system for speech recognition (see Holt, abstract);
Padmanabhan et al., US 6,219,638 B1 (hereinafter Padmanabhan), teaches telephone messaging and editing with speech to text, language translation, and text to speech synthesis features (see Padmanabhan, abstract and figure 1); and
Saindon et al., US 2002/0161578 A1 (hereinafter Saindon), teaches systems and methods for automated audio transcription, translation, and transfer (see Saindon, abstract), such as voice recognition software (see Saindon, ¶ 0011 and 0095), and using foreign language real-time translation software (see Saindon, ¶ 0016 and 0100).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel R Sellers whose telephone number is (571)272-7528. The examiner can normally be reached Mon - Fri 10:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Fan S Tsang can be reached at (571)272-7547. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Daniel R Sellers/Primary Examiner, Art Unit 2694