DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Claims Status
Claims 1-11 have been amended and are currently pending.
35 U.S.C. §112
Any and all interpretations under 35 U.S.C. §112(f) have been withdrawn due to amendment. Further, any associated 35 U.S.C. §112 rejections associated therewith are also withdrawn.
35 U.S.C. §101
The 35 U.S.C. §101 rejection is withdrawn due to amendment.
35 U.S.C. §102
The examiner thanks the applicant for attempting to further prosecution by amending the independent claims in an effort to separate itself from the applied prior art (Li).
It would seem applicant is alleging a ‘same feature value space’ is a temporal parameter insofar as that the argument is being made that compiling takes place after processing the frames and speech (Remarks, P. 7, V). After a thorough review of the instant disclosure, this does not appear to be the case. As far as can be inferred from the claim language, same space is to be construed as a location-based parameter (i.e., a moving mouth makes talking noises, a moving vehicle makes moving vehicle noises.).
If the implication is that there is a simultaneous processing and compilation of animation and sound, this is not reflected in the claimed invention and same is merely being conflated with simultaneous.
Therefore, the claimed invention in the base inventions needs to be amended to reflect such realities. The rejections under 35 U.S.C. §102 are maintained and presented below.
35 U.S.C. §103
As mentioned by applicant, the same deficiencies under the 35 U.S.C. §103 rejections rise and fall with the 35 U.S.C. §102 rejections above and therefore are maintained for the same reasons above.
The rejections under 35 U.S.C. §102 are maintained and presented below.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-2. 4-9 and 11 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al., (US Patent No. 11,417,041 B2) referred to as LI hereinafter.
Regarding claim 1, LI shows a voice processing device (FIG. 1, SERVER 120) comprising:
Circuitry (FIG. 7) that:
extracts a feature value of an avatar image (Col. 13, lines 15-55 disclose the multi-style landmark predictor, which has the ability to use unique feature values to alter or drive the animation and synchronization process.) from a same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); and
processes the voice uttered by the avatar image on a basis of the extracted feature value (FIG. 2, animation compiler 270 processes the audio in question and compiles it along with the created animation.).
Regarding claim 2, LI shows the limitations of claim 1 as applied above, and further shows wherein the circuitry extracts the feature value of the avatar image (Col. 9, lines 25-55, 'template facial landmarks.) by using a feature value extractor designed such that a feature value extracted from a voice and a feature value extracted from an avatar image created from a face image of a speaker who has uttered the voice are close feature values on the feature value space (Col. 9, lines 1-12 disclose wherein animation frames for an avatar are compiled along with input speech.)
or
a speaker feature value extractor designed such that a feature value extracted from a face image and a feature value extracted from an avatar image generated from the face image are close feature values on the space (Col. 9, lines 25-55 wherein 3D landmarks of an input image are used to warp the animation to synchronize with the input/animated voice.).
Regarding claim 4, LI shows the limitations of claim 2 as applied above, and further shows wherein the circuitry determines the feature value by using both the feature value extracted from the voice of the speaker and the feature value extracted from the avatar image (Col. 9, lines 25-55 wherein 3D landmarks of an input image are used to warp the animation to synchronize with the input/animated voice.).
Regarding claim 5, LI shows the limitations of claim 1 as applied above, and further shows wherein the circuitry extracts the feature value by using a feature extractor configured by a model learned by using a data set including a voice, a face image of a speaker who has uttered the voice, and an avatar image generated from the face image (Col. 10, lines 25-37 disclose a learning model to use as a baseline landmark from which to obtain features for synthesis.).
Regarding claim 6, LI shows the limitations of claim 1 as applied above, and further shows wherein the circuitry describes a voice, a face image of a speaker who has uttered the voice, and an avatar image generated from the face image by a common impression word and uses the common impression word as the feature value (Col. 13, lines 9-14 disclose landmark or 'impression' truths that a common and allow for comparison across various voice styles.).
Regarding claim 7, LI shows a voice processing method comprising: an extraction step of extracting, with circuitry, a feature value of an avatar image (Col. 13, lines 15-55 disclose the multi-style landmark predictor, which has the ability to use unique feature values to alter or drive the animation and synchronization process.) from the same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); and
Processing, with the circuitry, the voice uttered by the avatar image on a basis of the extracted feature value (FIG. 2, animation compiler 270 processes the audio in question and compiles it along with the created animation.).
Regarding claim 8, LI shows an information terminal (FIG. 1, SERVER 120) comprising:
Circuitry (FIG. 7) that:
inputs first data for creating an avatar image (FIG. 2, 225);
inputs second data for adjusting a voice of the avatar image (FIG. 2, 210); and
processes the voice of the avatar image on a basis of a feature value determined by using both a feature value extracted from the avatar image created on a basis of the first data and a feature value extracted from a voice of a speaker based on the second data (FIG. 2, 240-260), the feature value of the avatar image and the feature value of the voice of the speaker sharing a same feature value space (Col., 9, lines 1-15).
Regarding claim 9, LI shows an information processing device (FIG. 1, SERVER 120) comprising:
circuitry (FIG. 7) that:
causes a first model to extract a feature value of an avatar image (FIG. 2, 225) from a same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15);
causes a second model to convert a voice quality of the voice of the avatar image or performs voice synthesis on a basis of the feature value extracted by the first model (FIG. 2, 210); and
trains the first model and the second model by using a data set including at least two of a voice, a face image of a speaker who has uttered the voice, or an avatar image generated from the face image (FIG. 2, 240-260).
Regarding claim 11, LI shows a non-transitory computer-readable medium storing a computer program written in a computer-readable format that, when executed by a computer, cause the computer to perform a method comprising:
extracting a feature value of an avatar image (Col. 13, lines 15-55 disclose the multi-style landmark predictor, which has the ability to use unique feature values to alter or drive the animation and synchronization process.) from a same feature value space as a voice uttered by the avatar (Col. 9, lines 1-15); and
processing the voice uttered by the avatar image on a basis of the extracted feature value (FIG. 2, animation compiler 270 processes the audio in question and compiles it along with the created animation.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over LI in view of Shah et al., (US 2022/0070295 A1) referred to as SHAH hereinafter.
Regarding claim 3, LI shows the limitations of claim 2 as applied above, however failing to but SHAH does further show wherein the circuitry converts a voice quality of an input voice on a basis of the feature value in the feature value space or synthesizes a voice on a basis of the feature value in the feature value space (Fig. 5 shows the process of taking an agent's voice and converts it to a celebrity profile and synthesizes the voice for output.).
It is noted that both LE and SHAH are analogous to the claimed invention in that they both synthesize voice.
Therefore, it would have been obvious to one possessing ordinary skill in the art before the effective filing date of the claimed invention to modify LE in the spirit of SHAH because listening to the plain voice of an IVR or previously recorded messages can quickly become very boring. This may be particular true when the customer dislikes the voice. Once a customer is connected with a live agent, the customer may be more difficult to please if they have endured a lengthy session with a boring or unpleasant voice (SHAH, Paragraph [0004]).
Claim(s) 10 is rejected under 35 U.S.C. 103 as being unpatentable over LE in view of Port et al., (US 11,069,259 B2) referred to as PORT hereinafter.
Regarding claim 10, LI shows the limitations of claim 9 as applied above, however failing to but PORT does further show wherein the circuitry trains the first model and the second model by adversarial learning such that a discriminator that discriminates authenticity of a voice cannot discriminate the authenticity and a determiner that identifies a speaker of the voice cannot identify the speaker (Col. 4, lines 35-40 disclose adversarial learning models in order to identify discrimination and increase quality of output.).
It is noted that both LE and PORT are analogous to the claimed invention in that they both process audio voice signals.
Therefore, it would have been obvious to one possessing ordinary skill in the art before the effective filing date of the claimed invention to modify LE in the spirit of PORT because it ensures the sound sample [feature] utilized is most closely correlated with the desired impact (Col. 7, lines 10-13).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN W. RIDER whose telephone number is (571)270-1068. The examiner can normally be reached Monday-Friday, 7.00 am - 4.30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jamie J Atala can be reached at (571) 272-7384. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JUSTIN W. RIDER
Primary Patent Examiner
Art Unit 2486
/Justin W Rider/ Primary Patent Examiner, Art Unit 2486