Prosecution Insights
Last updated: August 17, 2026
Application No. 18/359,174

CONTEXT-AWARE VOICE SELF-AUTHORIZATION

Final Rejection §103§112
Filed
Jul 26, 2023
Examiner
WASHBURN, DANIEL C
Art Unit
2657
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
4 (Final)
50%
Grant Probability
Moderate
5-6
OA Rounds
1y 0m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
80 granted / 161 resolved
-12.3% vs TC avg
Strong +30% interview lift
Without
With
+29.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
7 currently pending
Career history
172
Total Applications
across all art units

Statute-Specific Performance

§101
12.2%
-27.8% vs TC avg
§103
53.0%
+13.0% vs TC avg
§102
15.4%
-24.6% vs TC avg
§112
11.3%
-28.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 161 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to the 35 U.S.C. 103 rejection of claim(s) 1-3, 6-10, 13-17, and 20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-3, 6-10, 13-17, and 20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 1, 8, and 15 describe, “generating a unique code, in numerical uniform time code format, for each item of contextual information”. However, the specification does not describe that each item of contextual information is in numerical uniform time code format. The specification, at ¶s [0043]-[0044], describes, “[0043] Next, at 206, the context-aware voice self-authorization program 150 captures contextual information for the captured user speech. Once a user is identified determined, the context-aware voice self-authorization program 150 may capture various items of contextual information surrounding the user speech, such as, but not limited to, current time, current date, location, recording/capturing device, sensor ID, and application ID.” “[0044] In one or more embodiments, the context-aware voice self-authorization program 150 may generate a unique code for each item of contextual information, such as a numerical uniform time code for the time at which the speech was initially captured or an internet protocol address for the sensor that captured the speech.” Thus, the specification provides support for the time contextual information being in numerical uniform time code format, but it does not provide support for all items of contextual information being in numerical uniform time code format, as currently claimed. The dependent claims are rejected due to their dependency on rejected base claims, as they inherit and do not correct the deficiency described above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 6-10, 13-17, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhimanaik et al. (US 10,079,024, herein “Bhimanaik”) in view of Wang et al. (CN-115035903-B, herein “Wang”), in view of Wouters et al. (US 2021/0050024, herein “Wouters”), in view of Reitz et al. (US 10,720,169, herein “Reitz”), in view of Chen (CN-107799121-A, herein “Chen”) in view of Chauhan (US 11,244,693, herein “Chauhan”) and further in view of Wang (CN-113807995-A, herein “Wang ‘995”). RE claims 1, 8, and 15, Bhimanaik describes a processor-implemented method and a computer system, the computer system comprising: one or more processors (FIG. 4 and col. 13 lns. 2-9 – processor 403), one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more tangible storage media for execution by at least one of the one or more processors via at least one of the one or more memories (FIG. 4 and col. 13 lns. 10-17 – memory 406, data store 212), wherein the computer system is capable of performing a method comprising: identifying a speaker in captured speech using two or more authentication techniques (FIG. 3B and col. 11 lns. 7-26, “In box 339, the voice-based authentication service 215 detects a voice authentication factor from the captured audio. For instance, the audio may contain a voice command for which authentication is required. In some cases, the voice-based authentication service 215 may cause a knowledge-based question to be asked via the speech synthesizer 248 (FIG. 2), where the user is prompted to supply a knowledge-based question answer 221 (FIG. 2). In box 342, the voice-based authentication service 215 determines whether the voice authentication factor in the audio matches the voice of the authorized user. In this regard, the voice-based authentication service 215 may perform an analysis of the voice embodied in the voice authentication factor, to include speed of delivery, pitch, spectral content, loudness, orientation relative to multiple microphones of the voice interface device 103, and/or other characteristics. These characteristics can then be compared with the voice profile 224 (FIG. 2) of the authorized user to determine a confidence score. The confidence score is then compared to a threshold to determine whether a match has occurred.”); capturing contextual information related to the captured speech (Col. 10 ln. 63 – Col. 11 ln. 6 “In box 333, the voice-based authentication service 215 detects a wake sound 218 (FIG. 2), which places the voice interface device 103 into an active listening mode. In box 336, the voice-based authentication service 215 configured a client computing device 206 (FIG. 2) to generate a current watermark signal for an authenticated user. In some cases, the client computing device 206 may be preconfigured with the information necessary to generate a watermark signal. For instance, the client computing device 206 may generate a watermark signal based at least in part on a seed and a current time and/or location.”); encoding the contextual information into [an audible or inaudible watermark] (Col. 11 lns. 1-6 “In some cases, the client computing device 206 may be preconfigured with the information necessary to generate a watermark signal. For instance, the client computing device 206 may generate a watermark signal based at least in part on a seed and a current time and/or location.” Col 5 lns. 3-11 “In various embodiments, the watermark generation rules 233 may specify that the watermark signal should be ultrasonic, or above the range of hearing for a human (e.g., greater than 20 kilohertz), or the watermark generation rules 233 may indicate that the watermark signal should be in an audible frequency range.”), embedding the inaudible [watermark] with the captured speech (Col. 11 ln. 62 – Col. 12 ln 3 “Otherwise, if the voice authentication factor is determined to match the authorized user's voice, the voice-based authentication service 215 continues from box 342 to box 348. In box 348, the voice-based authentication service 215 determines whether the captured audio includes the current watermark signal along with the voice authentication factor. In this regard, the voice-based authentication service 215 may perform an analysis on the captured audio for expected characteristics of the current watermark signal.”). Bhimanaik doesn’t describe encoding the contextual information into an audible voice, wherein the encoding further comprises: generating a unique code, in numerical uniform time code format, for each item of contextual information; appending the unique codes together in a numerical string, wherein an amount of unique codes appended is based on a security level preconfigured by a user, and wherein the unique codes are appended in a user-desired order of importance; and converting the numerical string into the audible voice using text-to-speech technology; modifying a speed of the audible voice so a length of the audible voice matches a length of the captured speech; [or] converting the audible voice to an inaudible, ultrasonic-ranged sound frequency; and embedding the inaudible sound frequency voice with the captured speech. However, Wang describes encoding the contextual information into an audible voice (Page 9: “As shown in FIG. 1, an injection method of a physical voice watermark by the embodiment of the present invention can include step S101-step S103: S101, determining a sound signal matched with the physical voice of the target scene, as a physical voice watermark signal;” Also see page 10: “In order to increase the unique characteristic of the sound signal, the sound signal matched with the physical speech of the target scene further can include the identification information of the target scene, a digital password, a character password and so on. Of course, the sound signal matched with the physical speech of the target scene can be not limited to the above contents. wherein the identification information of the target scene can be the location of the target scene, time, event and other content information, for example, can be a section of voice content, such as XX year X, XX company XX company and XX conference and so on. The digital password, character password can be a string of numbers generated randomly, character, also can be a string of preset number, character, the embodiment of the invention is not specifically limited.”), wherein the encoding further comprises: converting [a] numerical string into the audible voice using text-to-speech technology (see page 10: “In order to increase the unique characteristic of the sound signal, the sound signal matched with the physical speech of the target scene further can include the identification information of the target scene, a digital password, a character password and so on. Of course, the sound signal matched with the physical speech of the target scene can be not limited to the above contents. wherein the identification information of the target scene can be the location of the target scene, time, event and other content information, for example, can be a section of voice content, such as XX year X, XX company XX company and XX conference and so on. The digital password, character password can be a string of numbers generated randomly, character, also can be a string of preset number, character, the embodiment of the invention is not specifically limited.”); converting the audible voice to an inaudible, ultrasonic-ranged sound frequency (Page 10: “S102, according to the preset playing strategy, playing the physical voice watermark signal in the target scene, so that in the target scene in the presence of a recording device, the recording device recorded voice is overlapped voice; wherein the superposed voice is the voice after the voice superposition after playing the physical voice and the physical voice watermark signal;” Also see page 16: “In the above according to the preset playing strategy, before playing the physical voice watermark signal in the target scene, the embodiment of the invention claims a physical voice watermark injection method further comprises: S302, the physical voice watermark signal is modulated to the appointed frequency band, obtaining the physical voice watermark signal after modulation; the designated frequency band can be any frequency band, or a specific frequency band, such as human ear irperceptible sound wave frequency band, namely ultrasonic frequency band or infrasonic wave band.”); and embedding the inaudible sound frequency voice with the captured speech (Pages 11-12: “In this embodiment, because playing the physical voice watermark signal in the target scene, the physical voice watermark signal after playing the voice in the air interface with the physical voice of the target environment to overlap, which means the superposed after the superposed voice is embedded in the watermark information, at this time, if there is recording device in the target environment, then the recording device can only record the superposed voice, so that the recorded voice comprises a physical voice watermark signal, further, after playing the physical voice watermark signal, recording the appointed information of the physical voice watermark signal, so as to subsequently use the appointed information to trace source, The solution is that the voice of the embedded physical voice watermark signal is traced to provide the base.” Also see page 17: “In this embodiment, it can be recorded in the recording device of the voice embedded in the physical voice watermark signal, so as to provide the basis for the voice tracing. Further, the physical voice watermark signal is modulated to the specified frequency band, obtaining the physical voice watermark signal after modulation, further enriches the implementation way of the injection method of the physical voice watermark provided by the embodiment of the invention, and when the designated frequency band is an ultrasonic frequency band, It can not affect the normal conversation of the people in the target scene, injecting watermark for the sound signal of the physical world, or interference recording device for recording the voice.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Bhimanaik the system and method of encoding the contextual information into an audible voice; converting the audible voice to an inaudible, ultrasonic-ranged sound frequency; and embedding the inaudible sound frequency voice with the captured speech, as taught by Wang, as this would create a unique, environmentally-specific watermark that is then converted to an inaudible sound frequency range, which improves the security of the authentication process, as it relies on specific details associated with the environment to create a unique and traceable watermark. Bhimanaik in view of Wang doesn’t describe modifying a speed of the audible voice so a length of the audible voice matches a length of the captured speech. However, Wouters describes a system and method that includes modifying a speed of the [watermark] so a length of the [watermark] matches a length of the [synthetic] speech (See ¶ [0029]: “FIGS. 3A and 3B are schematic block diagrams illustrating an information content scaling processor 322 in an audio watermark processor 308, in accordance with an embodiment of the invention. Here, the scaling processor 322 is configured to vary an information content of the audio watermark signal based on at least one of an information content of the synthetic speech signal, a length of the synthetic speech signal, and a quality of the synthetic speech signal. For example, in FIG. 3A, upon determining that a synthetic speech signal, S.sub.1, 307a, received from (or being created by) a synthetic speech generator 306, has a high information content, long length and/or high quality, the scaling processor 322 of the audio watermark processor 308 scales the audio watermark, W.sub.1, accordingly. Thus, the audio watermarked synthetic speech signal, S.sub.1+W.sub.1, 309a, will be scaled by the scaling processor 322 to have a correspondingly high information content, long length and/or high quality, in such a situation.” Also see ¶ [0030]: “By contrast, in FIG. 3B, upon determining that a synthetic speech signal, S.sub.2, 307b, received from (or being created by) a synthetic speech generator 306, has a low information content, short length and/or low quality, the scaling processor 322 of the audio watermark processor 308 scales the audio watermark, W.sub.2, accordingly. Thus, the audio watermarked synthetic speech signal, S.sub.2+W.sub.2, 309b, will be scaled by the scaling processor 322 to have a correspondingly low information content, short length and/or low quality, in such a situation.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Bhimanaik in view of Wang a system and method that includes modifying a speed of the audible voice so a length of the audible voice matches a length of the captured speech, as suggested by Wouters, in order to improve security, by applying the audible voice to the entire captured speech signal, while also not increasing the length of the original captured speech signal, due to the audible voice signal being longer than the captured speech. Bhimanaik in view of Wang in view of Wouters doesn’t describe wherein the encoding further comprises: generating a unique code, in numerical uniform time code format, for each item of contextual information; appending the unique codes together in a numerical string, wherein an amount of unique codes appended is based on a security level preconfigured by a user, and wherein the unique codes are appended in a user-desired order of importance. However, Reitz describes a system and method wherein the encoding further comprises: generating a unique code, in numerical (Col. 6 lns. 27-34 “In the examples described in this section, the illustrative audio watermarks include tones that represent a recording device identifier (e.g., a serial number). Audio watermarks also may include tones that represent a timestamp (e.g., date, time), a user identifier (e.g., an identifier for a law enforcement officer associated with the recording device), location information (e.g., for GPS-equipped cameras) and/or other information.” Col. 6 lns. 60-66 “Together, strings of bits can represent integers, characters, or symbols (e.g., in an ASCII or UTF-8 format) that make up the recording device identifier or other information in the audio watermark. For example, the device identifier 45 can be transmitted with tones representing the bit string 0110100 (4 in ASCII format) followed by tones representing the bit string 0110101 (5 in ASCII format).”); [and] appending the unique codes together in a numerical string, (Col. 7 lns. 24-31 “In the examples described in this section, the nature of the information in the audio watermark can be determined from its expected length and its position. For example, an audio watermark signal containing a device identifier and a timestamp can take the form of: [start code] [device ID] [timestamp] [end code], where [start code] is 2 bytes, [device ID] is 4 bytes, [timestamp] is 8 bytes, and [end code] is 2 bytes.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Bhimanaik in view of Wang in view of Wouters a method and system wherein the encoding further comprises: generating a unique code, in numerical [and] appending the unique codes together in a numerical string, as taught by Reitz, in order to use a low bandwidth and easily identifiable technique of encoding contextual information into the watermark, which aids in accurately determining the source of the audio used in a replay attack. Bhimanaik in view of Wang in view of Wouters in view of Reitz doesn’t describe generating a unique code, in numerical uniform time code format. However, Chen describes a system and method that includes generating a unique code, in numerical uniform time code format (¶s [0060]-[0061]: “Step 105: Based on the preset generated heartbeat watermark data, the phase difference between the first maximum energy frequency and the second maximum energy frequency, adjust each audio frame in the second audio window and embed the heartbeat watermark data into each audio frame in the second audio window; the heartbeat watermark data is binary data composed of synchronization flags, Coordinated Universal Time and check data. [0060]” (emphasis added). “In this embodiment, the heartbeat watermark data consists of a 64-bit binary data string composed of a synchronization flag (2 bytes), a UTC time (4 bytes), and a verification data (2 bytes). When embedding a watermark, one bit of binary data is embedded in one audio frame, and 64 audio frames are enough to completely embed one watermark. [0061]”). (emphasis added). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Bhimanaik in view of Wang in view of Wouters in view of Reitz a system and method that includes generating a unique code, in numerical uniform time code format, as taught by Chen, in order to use a universally recognized format for the time data, which aids in accurately determining the source of the audio used in a replay attack. Bhimanaik in view of Wang in view of Wouters in view of Reitz in view of Chen doesn’t describe a system or method wherein an amount of unique codes appended is based on a security level preconfigured by a user. However, Chauhan describes a system and method wherein an amount of unique codes appended is based on a security level preconfigured by a user (col. 38 lns. 33-53: “The watermark engine 1114 can be configured to use metadata or information from the network application 1104, such as the network application's user's name and/or a confidentially status of the content, to generate the form and/or content of the watermark. The watermark engine 1114 can select and/or incorporate various information into the watermark, according to an identity or user profile of a user, the corresponding network application, or the type of content or information contained in the audio data stream 1112. The watermark engine 1114 can select a watermark from a plurality of existing or available watermarks (e.g., according to the type of information displayed, such as how sensitive or confidential the displayed information is, or assigned to particular network application(s)). In some embodiments, the watermark engine 1114 can select from one of a plurality available templates to be augmented, modified, customized or otherwise updated to form the watermark. The system can include a storage component accessible to the watermark engine 1114, for storing or maintaining the plurality of templates and/or available watermarks.” Also see col. 39 ln. 41 – col. 40 ln. 4: “The system 1100 can include one or more policies. The policies can be accessed, retrieved, selected, applied and/or otherwise used by the watermarking engine 1114 and/or the embedded browser 1108 to manage watermarks for network application(s) 1104 accessed via the embedded browser 1108. The policies can be stored or maintained in a storage component (e.g., the storage component that stores watermarks or related templates) on the client device 1106 and/or in a network location accessible by the watermark engine 1104. The watermarking engine 1114 can apply the one or more policies 1104 on a network application 1104 being accessed via the embedded browser 1108. For example, the watermarking engine 1114 can apply the one or more policies 1104 on metadata or a profile of the network application 1104 (e.g., that provides an indication of the type(s) of possible data/content involved, such as the sensitivity or confidentiality requirements of the data/content), and/or on information of the network application 1104 being displayed or rendered via the embedded browser 1108. For example, because the use of an embedded browser 1108 allows for real-time visibility into information provisioned or generated through the network application (accessed via the embedded browser 1108), for instance even before such information is rendered by the embedded browser 1108, the watermarking engine 1114 can analyze or otherwise process the information to determine and/or provide a suitable watermark (e.g., suitable contents for a watermark) to accompany the information being provided to a user. The watermarking engine 1114 can for instance apply one or more policies 1104 dynamically to the information to generate and/or dynamically update the watermark.” (emphasis added)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Bhimanaik in view of Wang in view of Wouters in view of Reitz in view of Chen a system and method wherein an amount of unique codes appended is based on a security level preconfigured by a user, as taught by Chauhan, in order to include more contextual information in situations where there is a high level of sensitivity of confidentiality, as more contextual information would aid in identifying the source of replay attacks that are attempting to access highly sensitive or confidential information. Bhimanaik in view of Wang in view of Wouters in view of Reitz in view of Chen in view of Chauhan doesn’t describe a system or method wherein the unique codes are appended in a user-desired order of importance. However, Wang ‘995 describes a system and method wherein the unique codes are appended in a user-desired order of importance (¶ [n0049]: “After obtaining at least one shooting information, a corresponding watermark can be determined for each shooting information. For example, geographic location information can determine a watermark that displays the geographic location, time information can determine a watermark that displays the shooting time, and subject category information can determine a watermark that displays the shooting category. The watermark content can be text, images, or a combination of text and images.” Also see ¶ [n0050]: “Each captured information can have a different priority. The priority can be preset by the electronic device, set according to user instructions, or determined based on the user's historical operating habits. Priority can be a priority value, with the order of priority determined by the size of the priority value, or it can be determined in other ways.” Also see ¶ [n0081]: “In this embodiment, when multiple shooting information is obtained, in order to determine the watermark pattern that matches the image to be processed based on multiple shooting information, the multiple shooting information can be combined. The combination can be a pairwise combination or a combination of all shooting information, and the corresponding watermark pattern can be determined by two or more combined shooting information. […] For example, the priority of combined information can be determined by superimposing the priority values corresponding to the shooting information in the combined information. For instance, the first combined information is composed of geographic location information and time information. The priority value of geographic location information is 10 and the priority value of time information is 8, so the priority value of the first combined information is 18. The second combined information is composed of geographic location information and theme category information. The priority value of theme category information is 12, so the priority value of the second combined information is 22. Therefore, the second combined information has a higher priority than the first combined information. The second combined information is used as the target combined information, and then the watermark pattern is determined based on the second combined information.” (emphasis added) Also see ¶ [n0083]: “In one embodiment, determining the watermark pattern based on the target combination information may include: determining the target watermark content corresponding to each shooting information in the target combination information; and arranging the target watermark content according to a preset layout to obtain the watermark pattern. In this embodiment, when multiple shooting information is obtained, in order to determine the watermark pattern that matches the image to be processed based on multiple shooting information, the target watermark content corresponding to each shooting information can be arranged in a preset manner so that the determined watermark pattern can display the content of each shooting information in the target combination information.” Finally, see ¶ [n0084]: “The preset arrangement can be a simple horizontal or vertical parallel arrangement, or a matrix arrangement or a curved arrangement. The target watermark content corresponding to each shooting information in the target combination information is arranged according to the preset arrangement to form the final watermark pattern. The preset layout can be preset by the electronic device or set by the user's instructions.” (emphasis added)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Bhimanaik in view of Wang in view of Wouters in view of Reitz in view of Chen in view of Chauhan a system and method wherein the unique codes are appended in a user-desired order of importance, as taught by Wang ‘995, in order to enable a user to specify which codes are the most critical to help in identifying the source of a replay attack, which aids in identifying the source of replay attacks, even when only a portion of the watermark data is received. RE claims 2, 9, and 16 Bhimanaik describes a method and system further comprising: receiving an audio file (Col. 10 lns. 51-58 “Beginning with box 330, the voice-based authentication service 215 receives audio captured from an audio input device 242 (FIG. 2) of a voice interface device 103 (FIG. 2). A client application 251 (FIG. 2) executed by the voice interface device 103 may encode the audio and stream the audio over the network 209 (FIG. 2) to the computing environment 203 for analysis by the voice-based authentication service 215.”); in response to determining the received audio file contains the inaudible sound frequency voice, extracting the embedded inaudible sound frequency voice (Col. 11 ln. 62 – Col. 12 ln. 8 “Otherwise, if the voice authentication factor is determined to match the authorized user's voice, the voice-based authentication service 215 continues from box 342 to box 348. In box 348, the voice-based authentication service 215 determines whether the captured audio includes the current watermark signal along with the voice authentication factor. In this regard, the voice-based authentication service 215 may perform an analysis on the captured audio for expected characteristics of the current watermark signal. For example, where the current watermark signal includes ultrasonic signals, the voice-based authentication service 215 may analyze the content of the captured audio to determine whether tones greater than 20 kilohertz are present, and then also the characteristics of such tones.”); and in response to determining the speaker should be authenticated, authenticating the speaker using the inaudible sound frequency voice (Col. 12 lns. 56-64 “If the captured audio does not include unexpected ambient audio, the voice-based authentication service 215 continues from box 351 to box 354. In box 354, the voice-based authentication service 215 approves authentication of the user. Consequently, an action requested by the user in the voice authentication factor, or subsequent to the voice authentication factor, may be approved.”). RE claims 3, 10, and 17, Bhimanaik describes a method and system wherein the embedding comprises overlaying the inaudible sound frequency voice over audio data so each plays simultaneously in a merged audio data file (Col. 11 ln. 62 – Col. 12 ln. 8 “Otherwise, if the voice authentication factor is determined to match the authorized user's voice, the voice-based authentication service 215 continues from box 342 to box 348. In box 348, the voice-based authentication service 215 determines whether the captured audio includes the current watermark signal along with the voice authentication factor. In this regard, the voice-based authentication service 215 may perform an analysis on the captured audio for expected characteristics of the current watermark signal. For example, where the current watermark signal includes ultrasonic signals, the voice-based authentication service 215 may analyze the content of the captured audio to determine whether tones greater than 20 kilohertz are present, and then also the characteristics of such tones.”). RE claims 6, 13, and 20, Bhimanaik describes a method and system wherein the two or more authentication techniques are selected from a group consisting of fingerprint scanning, voiceprint analysis, iris scanning, and password verification (Col. 11 lns. 7-27 “In box 339, the voice-based authentication service 215 detects a voice authentication factor from the captured audio. For instance, the audio may contain a voice command for which authentication is required. In some cases, the voice-based authentication service 215 may cause a knowledge-based question to be asked via the speech synthesizer 248 (FIG. 2), where the user is prompted to supply a knowledge-based question answer 221 (FIG. 2). In box 342, the voice-based authentication service 215 determines whether the voice authentication factor in the audio matches the voice of the authorized user. In this regard, the voice-based authentication service 215 may perform an analysis of the voice embodied in the voice authentication factor, to include speed of delivery, pitch, spectral content, loudness, orientation relative to multiple microphones of the voice interface device 103, and/or other characteristics. These characteristics can then be compared with the voice profile 224 (FIG. 2) of the authorized user to determine a confidence score. The confidence score is then compared to a threshold to determine whether a match has occurred.” Col. 11 lns. 35-41 “In some cases, a user may be prompted to provide other authentication factors in order to verify his or her identity, such as answering knowledge-based questions, providing another voice sample, providing a valid fingerprint, presenting a one-time password from a hardware token, obtaining corroboration from another authorized user, and so forth.”). RE claims 7 and 14, Bhimanaik describes a method and system further comprising: capturing audio data using a sensor communicatively coupled to a computing device (Col. 10 lns. 51-53 “Beginning with box 330, the voice-based authentication service 215 receives audio captured from an audio input device 242 (FIG. 2) of a voice interface device 103 (FIG. 2).” Col. 6 lns. 59-62 “The audio input devices 242 may comprise a microphone, a microphone-level audio input, a line-level audio input, or other types of input devices.”). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: see additional references cited on PTO-892. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel C Washburn whose telephone number is (571)272-5551. The examiner can normally be reached Monday-Friday 9:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 10 earlier events
Oct 09, 2025
Request for Continued Examination
Oct 13, 2025
Response after Non-Final Action
Jan 15, 2026
Non-Final Rejection mailed — §103, §112
Mar 13, 2026
Interview Requested
Mar 20, 2026
Applicant Interview (Telephonic)
Mar 20, 2026
Examiner Interview Summary
Mar 20, 2026
Response Filed
Jul 17, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12602555
METHOD FOR SEARCHING FOR TEXTS IN DIFFERENT LANGUAGES BASED ON PRONUNCIATION AND ELECTRONIC DEVICE APPLYING THE SAME
3y 3m to grant Granted Apr 14, 2026
Patent 12603084
METHOD, APPARATUS, AND COMPUTER-READABLE RECORDING MEDIUM FOR CONTROLLING RESPONSE UTTERANCE BEING REPRODUCED AND PREDICTING USER INTENTION
2y 7m to grant Granted Apr 14, 2026
Patent 12511480
Pattern Recognition Using NLP-Based Tokenizing and Clustering Models
2y 4m to grant Granted Dec 30, 2025
Patent 9614588
Smart Appliances
2y 2m to grant Granted Apr 04, 2017
Patent 8373711
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND COMPUTER-READABLE STORAGE MEDIUM
5y 2m to grant Granted Feb 12, 2013
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
50%
Grant Probability
80%
With Interview (+29.8%)
4y 1m (~1y 0m remaining)
Median Time to Grant
High
PTA Risk
Based on 161 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month