Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/01/26 has been entered.
This office action is in response to correspondence 06/01/26 regarding application 18/315,027, in which claims 1-7 and 15-20 were amended. Claims 1-7 and 15-20 are pending in the application and have been considered.
Response to Arguments
The examiner agrees with Applicant on page 8 that no new matter is added by the amendments to claims 1-7 and 15-20.
Applicant’s arguments on page 8 regarding the 35 U.S.C. 103 rejections based on Callaghan, Baldwin, Johnson, and Dutta have been considered but are moot in view of the new grounds for rejection based in part on the newly discovered reference to Yanagihara (US 20120166192), which discloses maintaining synchronization of speech and non-speech input received across various time periods on a mobile device text messaging application (see below).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Callaghan et al. (US 7406657) in view of Yanagihara (US 20120166192).
Consider claim 1, Callaghan discloses an apparatus, comprising:
a processor (a processor is inherent for the multithreaded processing of command events, Col 6 lines 52-58); and
a memory operably coupled to the processor, the memory storing instructions (a parser contained within the multimodal browser builds the representation from lines of code, Col 6 lines 22-28, for which memory coupled to the processor is inherent) to cause the processor to:
receive an indication of a start of a session, the session associated with a user and having an audio channel that is synchronized with a non-audio channel (user selects a field to fill within the form, which initiates audio presentation of the form field synchronized with display of the form field, Col 4 lines 30-48, and allows the user to fill in the form via verb interaction, Col 2 lines 20-23, by allowing navigation of the form and filling in of fields in response to prompts, Col 3 lines 47-53).
Callaghan does not specifically mention a session for at least one of a non-audio channel defined by a text messaging mobile application or a voice-powered automated assistant, the session having an audio channel over a first connection that is synchronized with a non-audio channel over a second connection different from the first connection; and after receipt of the indication of the start of the session and during a time period before receipt of a user prompt via the non- audio channel, maintain the audio channel over the first connection in a synchronized state with the non-audio channel over the second connection without causing audible output on the audio channel during the time period.
Yanagihara discloses a session for at least one of a non-audio channel defined by a text messaging mobile application or a voice-powered automated assistant (email application displayed on user interface with virtual keyboard, Fig. 1A, [0045]), the session having an audio channel over a first connection that is synchronized with a non-audio channel over a second connection different from the first connection (audio subsystem is coupled to speaker and microphone to facilitate voice-enabled functions, [0038], and touch screen controller is coupled to touch screen 346 for key entry on virtual keyboard, [0039], speech to text composition server synchronizing and combining the non-speech input with the speech data, [0058], Figs. 5-6); and after receipt of the indication of the start of the session and during a time period before receipt of a user prompt via the non-audio channel, maintain the audio channel over the first connection in a synchronized state with the non-audio channel over the second connection without causing audible output on the audio channel during the time period (after start streaming mode, a time when the user selects to enable speech input, and before receiving tap 2, RTP or RTSP system maintains synchronization of inputs received via speech and taps without causing audible output, Figs 5-6, [0031], [0058]-[0060]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan by including a session for at least one of a non-audio channel defined by a text messaging mobile application or a voice-powered automated assistant, the session having an audio channel over a first connection that is synchronized with a non-audio channel over a second connection different from the first connection; and after receipt of the indication of the start of the session and during a time period before receipt of a user prompt via the non- audio channel, maintain the audio channel over the first connection in a synchronized state with the non-audio channel over the second connection without causing audible output on the audio channel during the time period in order to avoid requiring the user to utilize escape sequences to exit and return to a speech input mode, as suggested by Yanagihara, predictably making speech recognition systems easier to use, as suggested by Yanagihara. The cited references are analogous art in the same field of multi-modal input.
Consider claim 15, Callaghan discloses a method, comprising:
receiving a representation of a request from a compute device associated with a user to complete a task (user selects a field to fill within the form using multi-modal browser, for which execution on a computer device is inherent, which initiates audio presentation of the form field synchronized with display of the form field, Col 4 lines 30-48);
causing an audio channel associated with the user to synchronize with at least one non-audio channel associated with the user (audio presentation of the form field is synchronized with display of the form field, Col 4 lines 30-48);
in response to receipt of the user prompt via the at least one non-audio channel (determining whether the user has selected field 104 via a touchscreen, Col 4 lines 30-41, or selected, “SUBMIT” via tactile input method, such as the touchscreen, Col 5 lines 12-18):
selecting a first audible output based on a determination that the user prompt is in accordance with the task (upon receiving “SUBMIT” via tactile input method, such as the touchscreen, Col 5 lines 12-18, the audio service thread progresses to the construct end element form the FORM element, and outputs message M from the audio element queue “This message is after the form”, Col 7 lines 9-13, Fig. 6);
selecting a second audible output based on a determination that the user prompt is not in accordance with the task (if the user has selected field 104 via a touchscreen, Col 4 lines 30-41, prompt H: “What is the customer’s problem” is output, Fig 6, by the audio element queue) and
sending a signal to cause one of the first audible output or the second audible output to be output on the audio channel (the thread doing the audio progression on behalf of the browser requests that progress be stopped, re-positions, and then restarted, to output the various prompts, Col 6-7 lines 29-13, Fig. 6).
Callaghan does not specifically mention a session for at least one of a non-audio channel defined by a text messaging mobile application or a voice-powered automated assistant, the session having an audio channel over a first connection that is synchronized with a non-audio channel over a second connection different from the first connection; and after receipt of the indication of the start of the session and during a time period before receipt of a user prompt via the non- audio channel, maintain the audio channel over the first connection in a synchronized state with the non-audio channel over the second connection without causing audible output on the audio channel during the time period.
Yanagihara discloses a session for at least one of a non-audio channel defined by a text messaging mobile application or a voice-powered automated assistant (email application displayed on user interface with virtual keyboard, Fig. 1A, [0045]), the session having an audio channel over a first connection that is synchronized with a non-audio channel over a second connection different from the first connection (audio subsystem is coupled to speaker and microphone to facilitate voice-enabled functions, [0038], and touch screen controller is coupled to touch screen 346 for key entry on virtual keyboard, [0039], speech to text composition server synchronizing and combining the non-speech input with the speech data, [0058], Figs. 5-6); and after receipt of the indication of the start of the session and during a time period before receipt of a user prompt via the non-audio channel, maintain the audio channel over the first connection in a synchronized state with the non-audio channel over the second connection without causing audible output on the audio channel during the time period (after start streaming mode, a time when the user selects to enable speech input, and before receiving tap 2, RTP or RTSP system maintains synchronization of inputs received via speech and taps without causing audible output, Figs 5-6, [0031], [0058]-[0060]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan by including a session for at least one of a non-audio channel defined by a text messaging mobile application or a voice-powered automated assistant, the session having an audio channel over a first connection that is synchronized with a non-audio channel over a second connection different from the first connection; and after receipt of the indication of the start of the session and during a time period before receipt of a user prompt via the non- audio channel, maintain the audio channel over the first connection in a synchronized state with the non-audio channel over the second connection without causing audible output on the audio channel during the time period for reasons similar to those for claim 1.
Consider claim 17, Callaghan discloses the prompt is a first prompt, the method further comprising: repeatedly determining whether a second prompt on the at least one non-audio channel has been received from the user (when the last field of the form is reached, a pause is generated and the system determines whether the user uses a tactile input method to either submit, cancel, or reset the form, Col 5 lines 3-18 before the pause expires, and if not, loops back to the top of the form, and upon reaching the last field, this solicitation of input for the user to submit the form is “a second prompt”, noting the claim does not require the content of the prompts to differ); sending a fourth signal to cause an inaudible output on the audio channel to the user in response to each determination that the second prompt on the at least one non-audio channel has not been received from the user (the audio progression loops back to the top of the form, i.e. remains synchronized with the displayed form, until the user uses a tactile input method to either submit, cancel, or reset the form, Col 5 lines 3-18, allowing the visual component of the form to remain synchronized with the audio presentation, Col 5 lines 34-40, and the electrical waveform output to the speaker is considered inaudible prior to being transduced into an acoustic pressure wave, for example, inherent in producing what is ultimately audio that is audible to the user from the WAV file, Col 6 lines 34-42); and in response to the determination that the second prompt on the at least one non-audio channel has been received from the user: selecting a fourth audible output based on an activity by the user on the at least one non-audio channel (if the user has selected field 104 via a touchscreen, Col 4 lines 30-41, outputting the next prompt in the audio element queue, i.e. “a fourth audible output”), and sending a fifth signal to cause the fourth audible output to be output on the audio channel (the thread doing the audio progression on behalf of the browser requests that progress be stopped, re-positions, and then restarted, to output the various prompts, the prompt after the third considered a “fourth audible output”, Col 6-7 lines 29-13, Fig. 6).
Claims 3, 4, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Callaghan et al. (US 7406657) in view of Yanagihara (US 20120166192), in further view of Johnson et al. (US 20030162561).
Consider claim 3, Callaghan discloses the memory further stores instructions to cause the processor to: in response to receipt of the user prompt via the non-audio channel, select an audible output based on an activity by the user on the non-audio channel (if the user has selected field 104 via a touchscreen, Col 4 lines 30-41, prompt H: “What is the customer’s problem” is output, Fig 6, by the audio element queue); select, at a first time, a first prompt from a plurality of prompts (a first WAV file is selected and output, Col 6 lines 34-42, at a first time, e.g. at “D” prompting “What is the customer name”, Fig 6); and select, at a second time after the first time, a second prompt from the plurality of prompt, wherein the audible output is selected based on the second prompt (a second WAV file is selected and output, Col 6 lines 34-42, at a second time, e.g. at “H” prompting “What is the customer problem?”, Fig 6), .
Callaghan and Yanagihara do not specifically mention select a first language and a second language.
Johnson discloses selecting a first language and a second language (voiceXML is the base language of the multimodal application, and the CMMT indicates a text mode containing the text in HTML, [0044] their “selections” implicit).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara by selecting a first language and a second language in order to allow concurrent input of information by a user in differing modes through differing applications, as suggested by Johnson, [0009], leading to predictable results of reducing the cumbersome need to manually switch modes during a session, as suggested by Johnson ([0006]). The cited references are analogous art in the field of multi-modal input.
Consider claim 4, Callaghan and Yanagihara do not, but Johnson discloses: the audio channel is associated with a first device type from a plurality of device types, the non-audio channel is associated with a second device type from the plurality of device types, the first device type includes a phone, a smart speaker, an earphone or an Internet of Things (IoT) device, the second device type includes a phone, a smart speaker, an earphone or an IoT device, and the first device type is different from the second device type (different modalities including voice and keyboard or touchscreen input on different devices, [0004], [0005], such as a cellular telephone, [0023], and speaker located on another device such as a PDA, [0024]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara such that the audio channel is associated with a first device type from a plurality of device types, the non-audio channel is associated with a second device type from the plurality of device types, the first device type includes a phone, a smart speaker, an earphone or an Internet of Things (IoT) device, the second device type includes a phone, a smart speaker, an earphone or an IoT device, and the first device type is different from the second device type for reasons similar to those for claim 3.
Consider claim 19, Callaghan and Yanagihara do not, but Johnson discloses: the audio channel is associated with a first device type from a plurality of device types, the non-audio channel is associated with a second device type from the plurality of device types, and the first device type includes a phone, a smart speaker, an earphone or an Internet of Things (IoT) device, the second device type includes a phone, a smart speaker, an earphone or an IoT device, and the first device type being different from the second device type (different modalities including voice and keyboard or touchscreen input on different devices, [0004], [0005], such as a cellular telephone, [0023], and speaker located on another device such as a PDA, [0024]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara such that the audio channel is associated with a first device type from a plurality of device types, the non-audio channel is associated with a second device type from the plurality of device types, and the first device type includes a phone, a smart speaker, an earphone or an Internet of Things (IoT) device, the second device type includes a phone, a smart speaker, an earphone or an IoT device, and the first device type being different from the second device type for reasons similar to those for claim 3.
Claims 5-7, 18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Callaghan et al. (US 7406657) in view of Yanagihara (US 20120166192), in further view of Dutta et al. (US 20200177730).
Consider claim 5, Callaghan discloses memory further stores instructions to cause the processor to: in response to receipt of the user prompt via the non-audio channel (determining whether the user has selected field 104 via a touchscreen, Col 4 lines 30-41, or selected, “SUBMIT” via tactile input method, such as the touchscreen, Col 5 lines 12-18), select an audible output based on an activity by the user on the non-audio channel, and receive a signal from a device for the non-audio channel, wherein the audible output is selected based on the signal from the device for the non-audio channel (if the user has selected field 104 via a touchscreen, Col 4 lines 30-41, prompt H: “What is the customer’s problem” is output, Fig 6, by the audio element queue, the thread doing the audio progression on behalf of the browser requests that progress be stopped, re-positions, and then restarted, to output the various prompts, Col 6-7 lines 29-13, Fig. 6).
Callaghan and Yanagihara do not specifically mention receive, via an application programming interface (API), a signal.
Dutta discloses receiving, via an application programming interface (API), a signal (an API call, [0067]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara by receiving, via an application programming interface (API), a signal in order to avoid disjointed communication, as suggested by Dutta ([0008]), predictably resulting in improved quality of experience for a customer, as suggested by Dutta ([0007]). The cited references are analogous art in the field of multi-modal input.
Consider claim 6, Callaghan discloses: during a first time period, the non-audio channel a first digital non-audio channel (determining that the user has selected field 104 via a touchscreen, Col 4 lines 30-41), and during a second time period after the first time period, the non-audio channel is a second digital non-audio channel different from the first digital non-audio channel, and the audible output is selected with respect to the second digital non-audio channel (after up to a 10 second delay, determining the user has selected, “SUBMIT” via tactile input method, such as the keyboard, Col 5 lines 12-18, Fig. 6).
Consider claim 7, Callaghan discloses the memory further stores instructions to cause the processor to:
after the start of the session and before the end of the session, perform at least one of: determine that a prompt on the audio channel received from the user includes an indication that the user would like to discontinue the non-audio channel (determining the user has spoken “CANCEL” command, Col 5 lines 17-19) or determine that the user prompt via the non-audio channel includes an indication that the user would like to discontinue the non-audio channel (determining the user has selected the “CANCEL” command via a tactile input method, Col 5 lines 17-19); terminate the non-audio channel of the session, in response to the indication that the user would like to discontinue the non-audio channel (the cycle continues until the user invokes a command to end, Col 7 liners 9-11, Step L, Fig. 6, which ends the audio/visual presentation, Col 6 lines 5-9);
Callaghan and Yanagihara do not specifically mention send, non-audio channel is terminated, a signal to connect a communication device of the user with a communication device of a live agent.
Dutta discloses sending a signal to connect a communication device of the user with a communication device of a live agent (connect to voice agent button 400, Fig. 4, [0062]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara by sending, after the non-audio channel is terminated as in Callaghan, a signal to connect a communication device of the user with a communication device of a live agent as disclosed by Dutta for reasons similar to those for claim 5.
Consider claim 18, Callaghan does not, but Yanagihara discloses the compute device is a mobile device (mobile device, [0004]); and automatically causing the audio channel associated with the user and over the first connection to synchronize with the at least one non-audio channel associated with the user and over the second connection (RTSP system automatically causes synchronization of inputs received via speech and taps without causing audible output, Figs 5-6, [0031], [0058]-[0060])
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan such that the compute device is a mobile device and automatically causing the audio channel associated with the user and over the first connection to synchronize with the at least one non-audio channel associated with the user and over the second connection for reasons similar to those for claim 1.
Callaghan and Yanagihara do not specifically mention transmitting a hyperlink to the mobile device via at least one of a text message or an email, the causing of the audio channel associated with the user to synchronize with the at least one non-audio channel associated with the user performed automatically in response to the user selecting the hyperlink.
Dutta discloses transmitting a hyperlink to the mobile device via at least one of a text message or an email, the causing of the audio channel associated with the user to synchronize with the at least one non-audio channel associated with the user performed automatically in response to the user selecting the hyperlink (the apparatus provides a message including a URL to the caller via email sent to a different device, which selection of triggers a linked web session, [0022]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara by transmitting a hyperlink to the mobile device via at least one of a text message or an email, the causing of the audio channel associated with the user to synchronize with the at least one non-audio channel associated with the user performed automatically in response to the user selecting the hyperlink for reasons similar to those for claim 5.
Consider claim 20, Callaghan and Yanagihara do not, but Dutta discloses the compute device is a first compute device, the method further comprising: causing a connection to a second compute device associated with at least one of a live chat or a live agent in response to an indication from the user to connect with at least one of the live chat or the live agent (connect to voice agent button 400, Fig. 4, [0062], which causes a connection from mobile phone of the user to the agent’s computing device).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Callaghan and Yanagihara such that the compute device is a first compute device, and by causing a connection to a second compute device associated with at least one of a live chat or a live agent in response to an indication from the user to connect with at least one of the live chat or the live agent for reasons similar to those for claim 5.
Allowable Subject Matter
Claims 2 and 16 are objected to as being dependent on a rejected base claim, but would be allowable if rewritten in independent form including all limitations for their respective parent claims.
The following is the examiner’s statement of reasons for indicating allowable subject matter:
Consider claim 2, the prior art does not disclose or suggest: “…the memory further stores instructions to cause the processor to cause an inaudible output on the audio channel to the user during the time period before receipt of the user prompt via the non-audio channel.”
Dimitroff et al. (US 20160301810) discloses receiving an indication of a start of a session associated with a user and having an audio channel that is synchronized with a non-audio channel (audio and video connections of communication devices, [0011], initiating a video call, [0017]); in response to a determination that the prompt on the non-audio channel has been received from the user (when the user selects a CD from the list, [0019], [0035]): selecting an audible output based on an activity by the user on the non-audio channel (the CD the user selected, [0019], [0035], and sending a signal to cause the audible output to be output on the audio channel (when the user selects a CD from the list, that CD is used to answer the video call, [0019], [0035]). Dimitroff further discloses sending inaudible outputs in response to a timer expiring while waiting for the user to select a CD to answer a call ([0019]-[0021]). However, there was no indication in the prior art to further modify Callaghan and Yanagihara with the inaudible output of Dimitroff by causing the processor to cause an inaudible output on the audio channel to the user during the time period before receipt of the user prompt via the non-audio channel.
Claim 16 recites similar limitations to those in claim 2, which are allowable over the prior art of record for similar reasons.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Jesse S Pullias/
Primary Examiner, Art Unit 2655 07/22/26