Prosecution Insights
Last updated: August 18, 2026
Application No. 18/487,419

AUDIO MANIPULATION OF EMULATED CONTENT

Non-Final OA §102§103
Filed
Oct 16, 2023
Examiner
BECKER, TYLER JUSTIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Motorola Mobility LLC
OA Round
3 (Non-Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
17 granted / 23 resolved
+11.9% vs TC avg
Moderate +6% lift
Without
With
+6.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
16 currently pending
Career history
45
Total Applications
across all art units

Statute-Specific Performance

§101
19.2%
-20.8% vs TC avg
§103
51.1%
+11.1% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
15.4%
-24.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 11th, 2026 has been entered. Response to Amendment The amendments filed May 11th, 2026 have been entered. Claims 1, 4, 9, and 15 have been amended. Claims 6 and 19 have been cancelled. Claims 21 and 22 have been added. Claims 1-5, 7-18, and 20-22 are pending and have been examined. Response to Arguments Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 3, 7, 21, and 22 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yang et al. (US Pat. Pub. No. 2017/0301329 A1 hereinafter Yang). Regarding claim 1, Yang discloses a media device, comprising: a memory configured to store original audio content (Yang, [0129]: "Embodiments in accordance with the present invention may take the form of, and/or be provided as, a computer program product encoded in a machine-readable medium as instruction sequences and other functional constructs of software, which may in turn be executed in a computational system (such as a iPhone handheld, mobile or portable computing device, or content server platform) to perform methods described herein. In general, a machine readable medium can include tangible articles that encode information in a form (e.g., as applications, source or object code, functionally descriptive information, etc.) readable by a machine (e.g., a computer, computational facilities of a mobile device or portable computing device, etc.) as well as tangible storage incident to transmission of the information."); and an audio manipulation manager implemented at least partially in hardware, the audio manipulation manager configured to: receive, from a user, input audio content that emulates the original audio content (Yang, [0071]: "As previously described as well as in the illustrated configuration, a user/vocalist sings along with a backing track karaoke style."); receive metadata associated with the original audio content, the metadata including a content creator voice category of the original audio content (Yang, [0073]: "Both pitch correction (to main or harmony pitches) and optionally added harmonies are chosen to correspond to a score 207, which in the illustrated configuration, is wirelessly communicated (261) to the device (e.g., from content server 110 to an iPhone handheld 101 or other portable computing device, recall FIG. 1) on which vocal capture and pitch-correction is to be performed, together with lyrics 208 and an audio encoding of the backing track 209."); determine a user voice category from the input audio content (Yang, [0050]: "Pitch detection and correction of a user's vocal performance are performed continuously and in real-time with respect to the audible rendering of the backing track at the handheld or portable computing device."); and transform the input audio content to manipulated audio content by changing the user voice category to the content creator voice category as the input audio content is received (Yang, [0071]: "Vocals captured (251) from a microphone input 201 are continuously pitch-corrected (252) to either main vocal pitch cues or, in some cases, to corresponding harmony cues in real-time for mix (253) with the backing track which is audibly rendered at one or more acoustic transducers 202. In some cases or embodiments, the audible rendering of captured vocals pitch corrected to “main” melody may optionally be mixed (254) with harmonies (HARMONY1, HARMONY2) synthesized from the captured vocals in accord with score coded offsets."). Regarding claim 3, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. Yang further discloses wherein the original audio content is a song, and the input audio content is a cover of the song (Yang, [0071]: "As previously described as well as in the illustrated configuration, a user/vocalist sings along with a backing track karaoke style."). Regarding claim 7, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. Yang further discloses wherein: a second content creator voice category is included with the metadata associated with the original audio content (Yang, [0053]: "Depending on the goals and implementation of a particular system, a user selectable vocal effects (EFX) schedule may include (in a computer readable media encoding) settings and/or parameters for one or more of spectral equalization, audio compression, pitch correction, stereo delay, and reverberation effects for application to one or more respective portions of the user's vocal performance. In some cases or embodiments, a vocal effects schedule may be characteristic of an artist, song or performance and may be applied to an audio encoding of the user's captured vocal performance to cause a derivative audio encoding or audible rendering to take on characteristics of the selected artist, song or performance."; [0057]: "In at least some cases or embodiments, the term vocal effects schedule may further encompass, an enumerated set of vocal EFX that varies in temporal or template correspondence with portions of a vocal score (e.g., with distinct vocal EFX sets for pre-chorus and chorus portions of a song and/or with distinct vocal effects sets for respective portions of a duet or other multi-vocalist performance)."); and the audio manipulation manager is configured to change the manipulated audio content from the content creator voice category to the second content creator voice category (Yang, [0058]: "Likewise, respective portions of a single vocal effects schedule (or for that matter, a pair of distinct vocal effects schedules) may be employed relative to respective vocal performance captures to provide appropriate and respective EFX for a vocal performance capture of a first portion of a duet performed by a first user and for a separate vocal performance capture of a second portion of a duet performed by a second user."). Regarding claim 21, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. Yang further discloses wherein the content creator voice category is received separately from the original audio content (Yang, [0073]: "Both pitch correction (to main or harmony pitches) and optionally added harmonies are chosen to correspond to a score 207, which in the illustrated configuration, is wirelessly communicated (261) to the device (e.g., from content server 110 to an iPhone handheld 101 or other portable computing device, recall FIG. 1) on which vocal capture and pitch-correction is to be performed, together with lyrics 208 and an audio encoding of the backing track 209."; Here, the content creator voice category and original audio content are seen as being received concurrently, but as separate data.). Regarding claim 22, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. Yang further discloses wherein the audio manipulation manager is configured to determine the user voice category from the input audio content as the input audio content is received (Yang, [0050]: "Pitch detection and correction of a user's vocal performance are performed continuously and in real-time with respect to the audible rendering of the backing track at the handheld or portable computing device."). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang as applied to claims 1, 3, 7, 21, and 22 above, and further in view of Shah et al. (US Pat. Pub. No. 2022/0070295 A1 hereinafter Shah). Regarding claim 2, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. However, Yang fails to expressly recite wherein to change the user voice category to the content creator voice category, the audio manipulation manager is configured to change a user tone and a user pitch to a content creator tone and a content creator pitch. Shah teaches wherein to change the user voice category to the content creator voice category, the audio manipulation manager is configured to change a user tone and a user pitch to a content creator tone and a content creator pitch (Shah, [0007]: "Using Artificial Intelligence (AI), such as a neural network, the voice (for audio calls or audio portion of an audio/video call) and face may be overlaid over the actual live agent's face and/or voice so the customer is presented with speech, and/or images, of the desired entity."; [0009]: "For speech, the tone and pitch of the voice of agent are also mapped to that of the celebrity's in real time."). Yang and Shah are analogous arts because they each belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang to incorporate the teachings of Shah to modify both pitch and tone of the user. This allows the user’s voice to better be modified to emulate a specific person’s voice (Shah, [0007]). By better emulating a specific voice, the systems final output is improved. Claim(s) 4 and 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang as applied to claims 1, 3, 7, 21, and 22 above, and further in view of Lindahl et al. (US Pat. Pub. No. 2013/0329908 A1 hereinafter Lindahl). Regarding claim 4, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. However, Yang fails to expressly recite wherein the audio manipulation manager is configured to detect that the media device is operating in an audio manipulation mode. Lindahl teaches wherein the audio manipulation manager is configured to detect that the media device is operating in an audio manipulation mode (Lindahl, [0009]: "To configure the audio beamforming settings, the computing system can detect a predetermined actively running application, such as a dictation application, a speech recognition application, an audio communications application, a video chat application, an audio recording application, or a music playback application."). Yang and Lindahl are analogous arts because they both belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang to incorporate the teachings of Lindahl to detect an audio manipulation mode. This allows the system to be configured based on a current state of a computing device (Lindahl, [0009]). This helps ensure the system works properly with other running programs. Regarding claim 8, the rejection of claim 1 is incorporated. Yang discloses all of the elements of the current invention as stated above. Yang further discloses wherein the audio manipulation manager is configured to: communicate the manipulated audio content to the audio output device for audio playback of the manipulated audio content (Yang, [0071]: "Vocals captured (251) from a microphone input 201 are continuously pitch-corrected (252) to either main vocal pitch cues or, in some cases, to corresponding harmony cues in real-time for mix (253) with the backing track which is audibly rendered at one or more acoustic transducers 202."). However, Yang fails to expressly recite wherein the audio manipulation manager is configured to: detect an audio output device is connected to the media device. Lindahl teaches wherein the audio manipulation manager is configured to: detect an audio output device is connected to the media device (Lindahl, [0009]: "in some cases, the system can detect at least one predetermined device setting, such as fan speed, current audio route, or a configuration of microphone and speaker placement."). The same motivation for claim 4 applies equally to claim 8. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang, in view of Lindahl, as applied to claims 4 and 8 above, and further in view of Wang et al. (US Pat. Pub. No. 2023/0281335 A1 hereinafter Wang). Regarding claim 5, the rejection of claim 4 is incorporated. Yang, in view of Lindahl, discloses all of the elements of the current invention as stated above. Lindahl further teaches wherein to detect that the media device is operating in the audio manipulation mode, the audio manipulation manager is configured to: detect an audio application is running in a foreground of the media device (Lindahl, [0009]: "To configure the audio beamforming settings, the computing system can detect a predetermined actively running application, such as a dictation application, a speech recognition application, an audio communications application, a video chat application, an audio recording application, or a music playback application."). The same motivation for claim 4 applies equally to claim 5. However, Yang, in view of Lindahl, fails to expressly recite wherein to detect that the media device is operating in the audio manipulation mode, the audio manipulation manager is configured to: detect the audio application is requesting use of a microphone of the media device. Wang teaches wherein to detect that the media device is operating in the audio manipulation mode, the audio manipulation manager is configured to: detect the audio application is requesting use of a microphone of the media device (Wang, [0054]: "the sound filter application can detect when a third party application is attempting to utilize a microphone, sound sensor, or the like to obtain information related to the user"). Yang, Lindahl, and Wang are analogous arts because they all belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang, as modified by the beamforming settings adjustment method of Lindahl, to incorporate the teachings of Wang to detect an audio application attempting to use a microphone. This ensures that no application is using a microphone without proper permissions (Wang, [0054]). This is important to protect the user’s privacy while using the system. Claim(s) 9, 11-13, 15, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang, in view of Metcalf. Regarding claim 9, Yang discloses a method, comprising: receiving, from a user, input audio content (Yang, [0071]: "As previously described as well as in the illustrated configuration, a user/vocalist sings along with a backing track karaoke style."); receiving metadata associated with the original audio content, the metadata including a content creator voice category from the original audio content (Yang, [0073]: "Both pitch correction (to main or harmony pitches) and optionally added harmonies are chosen to correspond to a score 207, which in the illustrated configuration, is wirelessly communicated (261) to the device (e.g., from content server 110 to an iPhone handheld 101 or other portable computing device, recall FIG. 1) on which vocal capture and pitch-correction is to be performed, together with lyrics 208 and an audio encoding of the backing track 209."); determining a user voice category from the input audio content (Yang, [0050]: "Pitch detection and correction of a user's vocal performance are performed continuously and in real-time with respect to the audible rendering of the backing track at the handheld or portable computing device."); and transforming the input audio content to manipulated audio content by changing the user voice category to the content creator voice category as the input audio content is received (Yang, [0071]: "Vocals captured (251) from a microphone input 201 are continuously pitch-corrected (252) to either main vocal pitch cues or, in some cases, to corresponding harmony cues in real-time for mix (253) with the backing track which is audibly rendered at one or more acoustic transducers 202. In some cases or embodiments, the audible rendering of captured vocals pitch corrected to “main” melody may optionally be mixed (254) with harmonies (HARMONY1, HARMONY2) synthesized from the captured vocals in accord with score coded offsets."). However, Yang fails to expressly recite identifying original audio content that is emulated by the input audio content. Metcalf teaches identifying original audio content that is emulated by the input audio content (Metcalf, [0016]: "the user sings or hums a portion of the song, in which case a pitch detector, a speech recognizer, or a combination of both is used in VUI platform 20 to identify the song."). Yang and Metcalf are analogous arts because they both belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang to incorporate the teachings of Metcalf to identify audio content. This allows the system to find a song quickly and little or no manual input from the user (Metcalf, [0004]). This improves the overall user experience of the system. Regarding claim 11, the rejection of claim 9 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. Yang further discloses wherein the original audio content is a song, and the input audio content is a cover of the song (Yang, [0071]: "As previously described as well as in the illustrated configuration, a user/vocalist sings along with a backing track karaoke style."). Regarding claim 12, the rejection of claim 9 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. Metcalf further teaches wherein identifying the original audio content includes at least one of: accessing an audio application to determine the original audio content; or detecting a tune of the input audio content (Metcalf, [0016]: "the user sings or hums a portion of the song, in which case a pitch detector, a speech recognizer, or a combination of both is used in VUI platform 20 to identify the song."). The same motivation for claim 9 applies equally to claim 12. Regarding claim 13, the rejection of claim 9 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. Yang further discloses wherein a second content creator voice category is included with the metadata associated with the original audio content (Yang, [0053]: "Depending on the goals and implementation of a particular system, a user selectable vocal effects (EFX) schedule may include (in a computer readable media encoding) settings and/or parameters for one or more of spectral equalization, audio compression, pitch correction, stereo delay, and reverberation effects for application to one or more respective portions of the user's vocal performance. In some cases or embodiments, a vocal effects schedule may be characteristic of an artist, song or performance and may be applied to an audio encoding of the user's captured vocal performance to cause a derivative audio encoding or audible rendering to take on characteristics of the selected artist, song or performance."; [0057]: "In at least some cases or embodiments, the term vocal effects schedule may further encompass, an enumerated set of vocal EFX that varies in temporal or template correspondence with portions of a vocal score (e.g., with distinct vocal EFX sets for pre-chorus and chorus portions of a song and/or with distinct vocal effects sets for respective portions of a duet or other multi-vocalist performance)."); and the transforming includes changing the manipulated audio content from the content creator voice category to the second content creator voice category (Yang, [0058]: "Likewise, respective portions of a single vocal effects schedule (or for that matter, a pair of distinct vocal effects schedules) may be employed relative to respective vocal performance captures to provide appropriate and respective EFX for a vocal performance capture of a first portion of a duet performed by a first user and for a separate vocal performance capture of a second portion of a duet performed by a second user."). Regarding claim 15, Yang discloses a system, comprising: a processor (Yang, [0039]: “In some embodiments in accordance with the present invention(s), a computer program product encoded in one or more non-transitory media, the computer program product includes instructions executable on a processor of the portable computing device to cause the portable computing device to perform the steps one of the above-described methods.”); and an audio manipulation manager implemented at least partially by the processor, configured to: detect an audio manipulation mode (Yang, [0071]: "As previously described as well as in the illustrated configuration, a user/vocalist sings along with a backing track karaoke style."); receive a content creator voice category associated with the original audio content (Yang, [0073]: "Both pitch correction (to main or harmony pitches) and optionally added harmonies are chosen to correspond to a score 207, which in the illustrated configuration, is wirelessly communicated (261) to the device (e.g., from content server 110 to an iPhone handheld 101 or other portable computing device, recall FIG. 1) on which vocal capture and pitch-correction is to be performed, together with lyrics 208 and an audio encoding of the backing track 209."); determine a user voice category from the input audio content (Yang, [0050]: "Pitch detection and correction of a user's vocal performance are performed continuously and in real-time with respect to the audible rendering of the backing track at the handheld or portable computing device."); and transform the input audio content to manipulated audio content by changing the user voice category to the content creator voice category as the input audio content is received (Yang, [0071]: "Vocals captured (251) from a microphone input 201 are continuously pitch-corrected (252) to either main vocal pitch cues or, in some cases, to corresponding harmony cues in real-time for mix (253) with the backing track which is audibly rendered at one or more acoustic transducers 202. In some cases or embodiments, the audible rendering of captured vocals pitch corrected to “main” melody may optionally be mixed (254) with harmonies (HARMONY1, HARMONY2) synthesized from the captured vocals in accord with score coded offsets."). However, Yang fails to expressly recite identify original audio content that is emulated by input audio content. Metcalf teaches identify original audio content that is emulated by input audio content (Metcalf, [0016]: "the user sings or hums a portion of the song, in which case a pitch detector, a speech recognizer, or a combination of both is used in VUI platform 20 to identify the song."). Yang and Metcalf are analogous arts because they both belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang to incorporate the teachings of Metcalf to identify audio content. This allows the system to find a song quickly and little or no manual input from the user (Metcalf, [0004]). This improves the overall user experience of the system. Regarding claim 18, the rejection of claim 15 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. Yang further discloses wherein the original audio content is a song, and the input audio content is a cover of the song (Yang, [0071]: "As previously described as well as in the illustrated configuration, a user/vocalist sings along with a backing track karaoke style."). Regarding claim 20, the rejection of claim 15 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. Yang further discloses wherein the audio manipulation manager is configured to: receive a second content creator voice category associated with the original audio content (Yang, [0053]: "Depending on the goals and implementation of a particular system, a user selectable vocal effects (EFX) schedule may include (in a computer readable media encoding) settings and/or parameters for one or more of spectral equalization, audio compression, pitch correction, stereo delay, and reverberation effects for application to one or more respective portions of the user's vocal performance. In some cases or embodiments, a vocal effects schedule may be characteristic of an artist, song or performance and may be applied to an audio encoding of the user's captured vocal performance to cause a derivative audio encoding or audible rendering to take on characteristics of the selected artist, song or performance."; [0057]: "In at least some cases or embodiments, the term vocal effects schedule may further encompass, an enumerated set of vocal EFX that varies in temporal or template correspondence with portions of a vocal score (e.g., with distinct vocal EFX sets for pre-chorus and chorus portions of a song and/or with distinct vocal effects sets for respective portions of a duet or other multi-vocalist performance)."); and change the manipulated audio content from the content creator voice category to the second content creator voice category (Yang, [0058]: "Likewise, respective portions of a single vocal effects schedule (or for that matter, a pair of distinct vocal effects schedules) may be employed relative to respective vocal performance captures to provide appropriate and respective EFX for a vocal performance capture of a first portion of a duet performed by a first user and for a separate vocal performance capture of a second portion of a duet performed by a second user."). Claim(s) 10 and 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang, in view of Metcalf, as applied to claims 9, 11-13, 15, and 18-20 above, and further in view of Shah. Regarding claim 10, the rejection of claim 9 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. However, Yang, in view of Metcalf, fails to expressly recite wherein changing the user voice category to the content creator voice category includes changing a user tone and a user pitch to a content creator tone and a content creator pitch. Shah teaches wherein changing the user voice category to the content creator voice category includes changing a user tone and a user pitch to a content creator tone and a content creator pitch (Shah, [0007]: "Using Artificial Intelligence (AI), such as a neural network, the voice (for audio calls or audio portion of an audio/video call) and face may be overlaid over the actual live agent's face and/or voice so the customer is presented with speech, and/or images, of the desired entity."; [0009]: "For speech, the tone and pitch of the voice of agent are also mapped to that of the celebrity's in real time."). Yang, Metcalf, and Shah are analogous arts because they all belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang, as modified by the music selection method of Metcalf, to incorporate the teachings of Shah to modify both pitch and tone of the user. This allows the user’s voice to better be modified to emulate a specific person’s voice (Shah, [0007]). By better emulating a specific voice, the systems final output is improved. Regarding claim 16, the rejection of claim 15 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. However, Yang, in view of Metcalf, fails to expressly recite wherein the user voice category includes a user tone and a user pitch, and the content creator voice category includes a content creator tone and a content creator pitch. Shah teaches wherein the user voice category includes a user tone and a user pitch, and the content creator voice category includes a content creator tone and a content creator pitch (Shah, [0007]: "Using Artificial Intelligence (AI), such as a neural network, the voice (for audio calls or audio portion of an audio/video call) and face may be overlaid over the actual live agent's face and/or voice so the customer is presented with speech, and/or images, of the desired entity."; [0009]: "For speech, the tone and pitch of the voice of agent are also mapped to that of the celebrity's in real time."). Yang, Metcalf, and Shah are analogous arts because they all belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang, as modified by the music selection method of Metcalf, to incorporate the teachings of Shah to modify both pitch and tone of the user. This allows the user’s voice to better be modified to emulate a specific person’s voice (Shah, [0007]). By better emulating a specific voice, the systems final output is improved. Regarding claim 17, the rejection of claim 16 is incorporated. Yang, in view of Metcalf and Shah, discloses all of the elements of the current invention as stated above. Shah further teaches wherein to change the user voice category to the content creator voice category, the audio manipulation manager is configured to change the user tone to the content creator tone and change the user pitch to the content creator pitch (Shah, [0007]: "Using Artificial Intelligence (AI), such as a neural network, the voice (for audio calls or audio portion of an audio/video call) and face may be overlaid over the actual live agent's face and/or voice so the customer is presented with speech, and/or images, of the desired entity."; [0009]: "For speech, the tone and pitch of the voice of agent are also mapped to that of the celebrity's in real time."). The same motivation for claim 16 applies equally to claim 17. Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang, in view of Metcalf, as applied to claims 9, 11-13, 15, and 18-20 above, and further in view of Lindahl. Regarding claim 14, the rejection of claim 9 is incorporated. Yang, in view of Metcalf, discloses all of the elements of the current invention as stated above. However, Yang, in view of Metcalf, fails to expressly recite detecting an audio manipulation mode by at least one of detecting an audio application is running, or detecting the audio application is requesting the input audio content. Lindahl teaches detecting an audio manipulation mode by at least one of detecting an audio application is running, or detecting the audio application is requesting the input audio content (Lindahl, [0009]: "To configure the audio beamforming settings, the computing system can detect a predetermined actively running application, such as a dictation application, a speech recognition application, an audio communications application, a video chat application, an audio recording application, or a music playback application."). Yang, Metcalf, and Lindahl are analogous arts because they all belong to the same field of audio processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the social music system of Yang, as modified by the music selection method of Metcalf, to incorporate the teachings of Lindahl to detect an audio manipulation mode. This allows the system to be configured based on a current state of a computing device (Lindahl, [0009]). This helps ensure the system works properly with other running programs. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER J BECKER whose telephone number is (703)756-1271. The examiner can normally be reached M-Th, 7:15am-5:45pm PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TYLER BECKER/ Examiner, Art Unit 2657 /DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 2 earlier events
Jan 12, 2026
Response Filed
Mar 10, 2026
Final Rejection mailed — §102, §103
Apr 30, 2026
Interview Requested
May 06, 2026
Applicant Interview (Telephonic)
May 06, 2026
Examiner Interview Summary
May 11, 2026
Request for Continued Examination
May 12, 2026
Response after Non-Final Action
Jun 03, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694228
REAL-TIME USER COMMUNICATION SENTIMENT DETECTION FOR DYNAMIC ANOMALY DETECTION AND MITIGATION
3y 4m to grant Granted Jul 28, 2026
Patent 12682113
SYSTEMS, METHODS, AND APPARATUSES FOR GENERATING STRUCTURED DATA FROM UNSTRUCTURED DATA USING NATURAL LANGUAGE PROCESSING TO GENERATE A SECURE MEDICAL DASHBOARD
3y 0m to grant Granted Jul 14, 2026
Patent 12651592
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR REAL-TIME LANGUAGE TRANSLATION USING GENERATIVE ARTIFICIAL INTELLIGENCE
3y 0m to grant Granted Jun 09, 2026
Patent 12632657
Joint Speech and Text Streaming Model for ASR
2y 10m to grant Granted May 19, 2026
Patent 12614560
REVERBERATION REMOVAL DEVICE, PARAMETER ESTIMATION DEVICE, REVERBERATION REMOVAL METHOD, PARAMETER ESTIMATION METHOD, AND PROGRAM
2y 9m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
80%
With Interview (+6.3%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month