Prosecution Insights
Last updated: August 17, 2026
Application No. 17/744,138

VOCAL RECORDING AND RE-CREATION

Final Rejection §103
Filed
May 13, 2022
Examiner
OPSASNICK, MICHAEL N
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Sony Group Corporation
OA Round
6 (Final)
82%
Grant Probability
Favorable
7-8
OA Rounds
0m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
753 granted / 919 resolved
+19.9% vs TC avg
Moderate +10% lift
Without
With
+10.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
34 currently pending
Career history
961
Total Applications
across all art units

Statute-Specific Performance

§101
19.5%
-20.5% vs TC avg
§103
33.5%
-6.5% vs TC avg
§102
30.0%
-10.0% vs TC avg
§112
5.3%
-34.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 919 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 19-37 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (GB 2571853) in view of Patel et al (20120089396). As per claim 1, Wang (GB 2571853) teaches a computer-implemented method comprising (page 5 – starting at line 6 “The implementation methods apdopted…”): recording, by a first device, audio of an utterance of a first user that is speaking to a second user of a second device that is located over a network from the first device (as, communicating from one virtual reality device to another virtual reality device – pp5, starting at line 7 “Virtual reality platform receives …from the first virtual reality device…second virtual reality devce”; and as acquiring user’s speech signal by the speech acquisition module – middle of pp 6, “The mobile voice terminal….”; converting the utterance recorded in the audio to text (as, preprocessing the speech signals into text – middle of pp 6, “speech recognition module aimed at converting….text message..”); determining, by the first device and based on an analysis of the audio, one or more derived speech characteristics of the first user ((as, processing and operating, on the speech characteristics – pp7, line 4 “recognition module consists of speech characteristic extraction…”; deriving emotion and expressions from the input speech – middle pp 6, “extraction module of speech’s emotional characteristic parameters intended to extract the parameters with emotional characteristics in pre-process speech signal…”); determining, by the first device, one or more preferred speech characteristics that are predefined by the first user for synthesizing speech outputs at devices that are located over the network; generating, by the first device, metadata generating, by the first device, data packets based at least on compressing the text and the metadata that specifies the transmitting, by the first device, the data packets over the network to the second device (as, re-creating the audio by the action match unit of the virtual environment terminal matches the emotional characteristic received – pp 7 last 6, to pp 8 line 2). Although Wang (GB 2571853) teaches the concept of generating a persona (virtual reality character) and automating the character to express the emotions derived from the user’s input speech (see bottom of pp 7, “virtual character;s emotional expressions and actions”) and further contemplates the intonation and speed of the speech emotion in the played speech message – pp 8, lines 1-6; Wang (GB 2571853) does not explicitly teach further details of the speech signal processing tied into a persona of the user. Patel et al (20120089396) teaches specific focus on speech parameter processing based on the language type (para 0130, processing acoustic cues at differing frequency/tones tied to the emotion); tuning the characteristics based on the desired results (preferences set by choosing an emotional category closest to the sample – see para 0121, and choosing/selecting the percentage that is ‘close enough’ – end of para 0121); speech characteristics being any one of speech, pitch, spacing, volume, etc. tied to the emotions and verbal expressions – para 0030, disclosing fundamental frequency, pitch, intensity, loudness, speaker rate, etc. etc. Therefore, it would have been obvious to one of ordinary skill in the art of emotion detection/extraction, from speech information, to enhance the system of Wang (GB 2571853) with the further processing of speech characteristics as taught by Patel et al (20120089396) because it would advantageously improve upon the end-user understanding/ perceptions (see Patel et al (20120089396), para 0102 – see perceptual improvements in the listed categories). As per claim 19, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, wherein the one or more of the preferred speech characteristics specify a preferred tone, volume, pacing or pitch of synthesized speech outputs (see Wang (GB 2571853) , pp 8, top, inferring intonation/speed of speech; and Patel et al (20120089396) -- , para 0030, disclosing fundamental frequency, pitch, intensity, loudness, speaker rate, etc.). As per claim 20, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, wherein one or more of the derived speech characteristics specify a discerned emotion or intent (see Wang (GB 2571853), deriving emotion and expressions from the input speech – middle pp 6, “extraction module of speech’s emotional characteristic parameters intended to extract the parameters with emotional characteristics in pre-process speech signal…” ). As per claims 21-23,25, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the first device comprises a gaming console; transmitting the data packets over the network; the audio of the utterance of the first user is recorded via a player chat application that is implemented by the first device (see Wang (GB 2571853), pp6, wherein the environment is a mobile voice terminal tied to a virtual environment terminal (e.g, gaming console – claim 21) tied to a network for a 3D virtual environment for multiple users – abstract (claim 22,23); also utilizing video footage – pp4, second full paragraph – “video device” recording and displaying – see following paragraph “Specific Operations:…”). As per claim 24, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, wherein the one or more derived speech characteristics comprise a discerned background noise that is associated with the audio (as, monitoring and processing, for noise in the signal – see Patel et al (20120089396), para 0062 – 0063). As per claim 26, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, comprising receiving, by the first device, data indicating one or more values specified by one or more respective tunable knobs that were adjusted by the first user at the first device for predefining the one or more respective preferred speech characteristics for synthesizing speech outputs (see Wang (GB 2571853), allowing the operator to make adjustments on text/voice/video – pp 14, bottom “Upon selection…”, reflecting back to “Upon recording…text,voice, video; in view of the mapping to claim 1 above, the adjustments would follow through to the speech characteristics, in the combination of Wang (GB 2571853) in view of Patel et al (20120089396)). Claims 27-35 are system claims that perform the commonly shared steps of method claims 1, 19-26 above and as such, claims 27-35 are similar in scope and content to claims 1,19-26; therefore, claims 27-35 are rejected under similar rationale as presented against claims 1,19-26 above. Furthermore, Wang (GB 2571853) teaches processors/storage devices, for executable step retrieval and processing (see Wang (GB 2571853), pp 6, middle – processor/memory. Claims 36, 37 are non-transitory computer readable media claims whose instructions/steps are executed by a processor with said steps, are found throughout in method claims 1,19-26 above and as such, claims 36,37 are similar in scope and content to claims 1,19-26; therefore, claims 36,37 are rejected under similar rationale as presented against claims 1,19-26 above. Furthermore, Wang (GB 2571853) teaches processors/storage devices, for executable step retrieval and processing (see Wang (GB 2571853), pp 6, middle – processor/memory. Response to Arguments Applicant’s arguments with respect to claim(s) have been considered but are moot because the new ground of rejection refer to new citations/combinations not previously presented. Furthermore, examiner notes that applicants arguments are toward the amended claim language; examiner points to the further detailed explanations/mappings to the Wang (GB 2571853)/ Patel et al (20120089396) references. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please note the reference cited on the PTO-892 form. In further detail, examiner notes the following references pertaining to applicants spec/claim scope: Mennicken et al (20210104220) teaches the selection of a persona from a plurality of personas based on the user -- Mennicken et al (20210104220) – see para 0082, wherein the user-selected characteristics to be used for specific users – end of para 0082. Zeng et al (20210073526) teaches extraction of emotional information, from the information of the user (in fact, Zeng et al teaches extraction of emotional information from multiple modes – visual, audio, and text – see para 0053; and further, are compared for verification through a fusing process based on semantic meaning and maximized offset – para 0053, and in further detail para 0054, and the like). Iwase et al (20200051545) teaches user selectable tones/types – para 0081 Yamagami et al (20090259475) teaches user selectable voice quality changes – para 0162 Sohn et al (20140022370) teaches storing of facial-emotion relationships for predicting from audio – para 0016 Socolof et al (20210097468) teaches analysis of sentiment clusters during a user interaction (para 0035, 0078). Ahn et al (20140093849) teaches analysis and selection of estimated emotions from a collection of emotion vectors (para 0006). Kang et al (20100121804) teaches estimating emotion using emotion vector parameters for comparison (para 0019). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Michael N Opsasnick/Primary Examiner, Art Unit 2658 07/19/2026
Read full office action

Prosecution Timeline

Show 9 earlier events
Apr 10, 2025
Examiner Interview Summary
May 09, 2025
Response Filed
Aug 08, 2025
Final Rejection mailed — §103
Jan 08, 2026
Request for Continued Examination
Jan 23, 2026
Response after Non-Final Action
Jan 27, 2026
Non-Final Rejection mailed — §103
Apr 27, 2026
Response Filed
Jul 22, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707212
SYSTEM FOR FACILITATING IN-PERSON INTERACTION BETWEEN MULTIUSER VIRTUAL ENVIRONMENT USERS WHOSE AVATARS HAVE INTERACTED VIRTUALLY
4y 4m to grant Granted Aug 11, 2026
Patent 12699201
Audio rendering of an electromagnetic metal detection signal
3y 9m to grant Granted Aug 04, 2026
Patent 12676164
System and Method for Modulation Domain-Based Audio Signal Encoding
3y 0m to grant Granted Jul 07, 2026
Patent 12658172
COMPUTING SYSTEM FOR UNSUPERVISED EMOTIONAL TEXT TO SPEECH TRAINING
4y 3m to grant Granted Jun 16, 2026
Patent 12651607
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
2y 2m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

7-8
Expected OA Rounds
82%
Grant Probability
92%
With Interview (+10.1%)
3y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 919 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month