Prosecution Insights
Last updated: August 18, 2026
Application No. 18/837,211

VOICE PROCESSING DEVICE, VOICE PROCESSING METHOD, INFORMATION TERMINAL, INFORMATION PROCESSING DEVICE, AND COMPUTER PROGRAM

Final Rejection §102§103§112
Filed
Aug 09, 2024
Priority
Mar 04, 2022 — JP 2022-033951 +1 more
Examiner
RIDER, JUSTIN W
Art Unit
2486
Tech Center
2400 — Computer Networks
Assignee
Sony Group Corporation
OA Round
2 (Final)
84%
Grant Probability
Favorable
3-4
OA Rounds
1y 5m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
220 granted / 262 resolved
+26.0% vs TC avg
Moderate +12% lift
Without
With
+12.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
19 currently pending
Career history
283
Total Applications
across all art units

Statute-Specific Performance

§101
15.2%
-24.8% vs TC avg
§103
38.5%
-1.5% vs TC avg
§102
33.5%
-6.5% vs TC avg
§112
7.3%
-32.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 262 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Claims Status Claims 1-11 have been amended and are currently pending. 35 U.S.C. §112 Any and all interpretations under 35 U.S.C. §112(f) have been withdrawn due to amendment. Further, any associated 35 U.S.C. §112 rejections associated therewith are also withdrawn. 35 U.S.C. §101 The 35 U.S.C. §101 rejection is withdrawn due to amendment. 35 U.S.C. §102 The examiner thanks the applicant for attempting to further prosecution by amending the independent claims in an effort to separate itself from the applied prior art (Li). It would seem applicant is alleging a ‘same feature value space’ is a temporal parameter insofar as that the argument is being made that compiling takes place after processing the frames and speech (Remarks, P. 7, V). After a thorough review of the instant disclosure, this does not appear to be the case. As far as can be inferred from the claim language, same space is to be construed as a location-based parameter (i.e., a moving mouth makes talking noises, a moving vehicle makes moving vehicle noises.). If the implication is that there is a simultaneous processing and compilation of animation and sound, this is not reflected in the claimed invention and same is merely being conflated with simultaneous. Therefore, the claimed invention in the base inventions needs to be amended to reflect such realities. The rejections under 35 U.S.C. §102 are maintained and presented below. 35 U.S.C. §103 As mentioned by applicant, the same deficiencies under the 35 U.S.C. §103 rejections rise and fall with the 35 U.S.C. §102 rejections above and therefore are maintained for the same reasons above. The rejections under 35 U.S.C. §102 are maintained and presented below. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-2. 4-9 and 11 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al., (US Patent No. 11,417,041 B2) referred to as LI hereinafter. Regarding claim 1, LI shows a voice processing device (FIG. 1, SERVER 120) comprising: Circuitry (FIG. 7) that: extracts a feature value of an avatar image (Col. 13, lines 15-55 disclose the multi-style landmark predictor, which has the ability to use unique feature values to alter or drive the animation and synchronization process.) from a same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); and processes the voice uttered by the avatar image on a basis of the extracted feature value (FIG. 2, animation compiler 270 processes the audio in question and compiles it along with the created animation.). Regarding claim 2, LI shows the limitations of claim 1 as applied above, and further shows wherein the circuitry extracts the feature value of the avatar image (Col. 9, lines 25-55, 'template facial landmarks.) by using a feature value extractor designed such that a feature value extracted from a voice and a feature value extracted from an avatar image created from a face image of a speaker who has uttered the voice are close feature values on the feature value space (Col. 9, lines 1-12 disclose wherein animation frames for an avatar are compiled along with input speech.) or a speaker feature value extractor designed such that a feature value extracted from a face image and a feature value extracted from an avatar image generated from the face image are close feature values on the space (Col. 9, lines 25-55 wherein 3D landmarks of an input image are used to warp the animation to synchronize with the input/animated voice.). Regarding claim 4, LI shows the limitations of claim 2 as applied above, and further shows wherein the circuitry determines the feature value by using both the feature value extracted from the voice of the speaker and the feature value extracted from the avatar image (Col. 9, lines 25-55 wherein 3D landmarks of an input image are used to warp the animation to synchronize with the input/animated voice.). Regarding claim 5, LI shows the limitations of claim 1 as applied above, and further shows wherein the circuitry extracts the feature value by using a feature extractor configured by a model learned by using a data set including a voice, a face image of a speaker who has uttered the voice, and an avatar image generated from the face image (Col. 10, lines 25-37 disclose a learning model to use as a baseline landmark from which to obtain features for synthesis.). Regarding claim 6, LI shows the limitations of claim 1 as applied above, and further shows wherein the circuitry describes a voice, a face image of a speaker who has uttered the voice, and an avatar image generated from the face image by a common impression word and uses the common impression word as the feature value (Col. 13, lines 9-14 disclose landmark or 'impression' truths that a common and allow for comparison across various voice styles.). Regarding claim 7, LI shows a voice processing method comprising: an extraction step of extracting, with circuitry, a feature value of an avatar image (Col. 13, lines 15-55 disclose the multi-style landmark predictor, which has the ability to use unique feature values to alter or drive the animation and synchronization process.) from the same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); and Processing, with the circuitry, the voice uttered by the avatar image on a basis of the extracted feature value (FIG. 2, animation compiler 270 processes the audio in question and compiles it along with the created animation.). Regarding claim 8, LI shows an information terminal (FIG. 1, SERVER 120) comprising: Circuitry (FIG. 7) that: inputs first data for creating an avatar image (FIG. 2, 225); inputs second data for adjusting a voice of the avatar image (FIG. 2, 210); and processes the voice of the avatar image on a basis of a feature value determined by using both a feature value extracted from the avatar image created on a basis of the first data and a feature value extracted from a voice of a speaker based on the second data (FIG. 2, 240-260), the feature value of the avatar image and the feature value of the voice of the speaker sharing a same feature value space (Col., 9, lines 1-15). Regarding claim 9, LI shows an information processing device (FIG. 1, SERVER 120) comprising: circuitry (FIG. 7) that: causes a first model to extract a feature value of an avatar image (FIG. 2, 225) from a same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); causes a second model to convert a voice quality of the voice of the avatar image or performs voice synthesis on a basis of the feature value extracted by the first model (FIG. 2, 210); and trains the first model and the second model by using a data set including at least two of a voice, a face image of a speaker who has uttered the voice, or an avatar image generated from the face image (FIG. 2, 240-260). Regarding claim 11, LI shows a non-transitory computer-readable medium storing a computer program written in a computer-readable format that, when executed by a computer, cause the computer to perform a method comprising: extracting a feature value of an avatar image (Col. 13, lines 15-55 disclose the multi-style landmark predictor, which has the ability to use unique feature values to alter or drive the animation and synchronization process.) from a same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); and processing the voice uttered by the avatar image on a basis of the extracted feature value (FIG. 2, animation compiler 270 processes the audio in question and compiles it along with the created animation.). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over LI in view of Shah et al., (US 2022/0070295 A1) referred to as SHAH hereinafter. Regarding claim 3, LI shows the limitations of claim 2 as applied above, however failing to but SHAH does further show wherein the circuitry converts a voice quality of an input voice on a basis of the feature value in the feature value space or synthesizes a voice on a basis of the feature value in the feature value space (Fig. 5 shows the process of taking an agent's voice and converts it to a celebrity profile and synthesizes the voice for output.). It is noted that both LE and SHAH are analogous to the claimed invention in that they both synthesize voice. Therefore, it would have been obvious to one possessing ordinary skill in the art before the effective filing date of the claimed invention to modify LE in the spirit of SHAH because listening to the plain voice of an IVR or previously recorded messages can quickly become very boring. This may be particular true when the customer dislikes the voice. Once a customer is connected with a live agent, the customer may be more difficult to please if they have endured a lengthy session with a boring or unpleasant voice (SHAH, Paragraph [0004]). Claim(s) 10 is rejected under 35 U.S.C. 103 as being unpatentable over LE in view of Port et al., (US 11,069,259 B2) referred to as PORT hereinafter. Regarding claim 10, LI shows the limitations of claim 9 as applied above, however failing to but PORT does further show wherein the circuitry trains the first model and the second model by adversarial learning such that a discriminator that discriminates authenticity of a voice cannot discriminate the authenticity and a determiner that identifies a speaker of the voice cannot identify the speaker (Col. 4, lines 35-40 disclose adversarial learning models in order to identify discrimination and increase quality of output.). It is noted that both LE and PORT are analogous to the claimed invention in that they both process audio voice signals. Therefore, it would have been obvious to one possessing ordinary skill in the art before the effective filing date of the claimed invention to modify LE in the spirit of PORT because it ensures the sound sample [feature] utilized is most closely correlated with the desired impact (Col. 7, lines 10-13). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN W. RIDER whose telephone number is (571)270-1068. The examiner can normally be reached Monday-Friday, 7.00 am - 4.30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jamie J Atala can be reached at (571) 272-7384. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. JUSTIN W. RIDER Primary Patent Examiner Art Unit 2486 /Justin W Rider/ Primary Patent Examiner, Art Unit 2486
Read full office action

Prosecution Timeline

Aug 09, 2024
Application Filed
Feb 05, 2026
Non-Final Rejection mailed — §102, §103, §112
May 13, 2026
Response Filed
Jul 16, 2026
Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707064
VIDEO ENCODING/DECODING METHOD AND APPARATUS APPLYING A MERGE MODE WITH MOTION VECTOR DIFFERENCE TO COMBINED AN INTER/INTRA PREDICTION MODE OR A GEOMETRIC PARTITIONING MODE
2y 4m to grant Granted Aug 11, 2026
Patent 12695910
VIDEO ENCODING, DECODING METHOD AND DECODER
1y 9m to grant Granted Jul 28, 2026
Patent 12695866
INTRA MODE CODING BASED ON TEMPLATE
1y 9m to grant Granted Jul 28, 2026
Patent 12695922
ENTROPY DECODING METHOD, AND DECODING APPARATUS USING SAME
1y 5m to grant Granted Jul 28, 2026
Patent 12688640
ANIMATED DECORATIVE ITEMS WITH SENSOR INPUTS
3y 1m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
84%
Grant Probability
96%
With Interview (+12.3%)
3y 5m (~1y 5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 262 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month