Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 19-37 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (GB 2571853) in view of Patel et al (20120089396).
As per claim 1, Wang (GB 2571853) teaches a computer-implemented method comprising (page 5 – starting at line 6 “The implementation methods apdopted…”):
recording, by a first device, audio of an utterance of a first user that is speaking to a second user of a second device that is located over a network from the first device (as, communicating from one virtual reality device to another virtual reality device – pp5, starting at line 7 “Virtual reality platform receives …from the first virtual reality device…second virtual reality devce”; and as acquiring user’s speech signal by the speech acquisition module – middle of pp 6, “The mobile voice terminal….”;
converting the utterance recorded in the audio to text (as, preprocessing the speech signals into text – middle of pp 6, “speech recognition module aimed at converting….text message..”);
determining, by the first device and based on an analysis of the audio, one or more derived speech characteristics of the first user ((as, processing and operating, on the speech characteristics – pp7, line 4 “recognition module consists of speech characteristic extraction…”; deriving emotion and expressions from the input speech – middle pp 6, “extraction module of speech’s emotional characteristic parameters intended to extract the parameters with emotional characteristics in pre-process speech signal…”);
determining, by the first device, one or more preferred speech characteristics that are
predefined by the first user for synthesizing speech outputs at devices that are located over the
network;
generating, by the first device, metadata
generating, by the first device, data packets based at least on compressing the text and
the metadata that specifies the
transmitting, by the first device, the data packets over the network to the second device (as, re-creating the audio by the action match unit of the virtual environment terminal matches the emotional characteristic received – pp 7 last 6, to pp 8 line 2).
Although Wang (GB 2571853) teaches the concept of generating a persona (virtual reality character) and automating the character to express the emotions derived from the user’s input speech (see bottom of pp 7, “virtual character;s emotional expressions and actions”) and further contemplates the intonation and speed of the speech emotion in the played speech message – pp 8, lines 1-6; Wang (GB 2571853) does not explicitly teach further details of the speech signal processing tied into a persona of the user. Patel et al (20120089396) teaches specific focus on speech parameter processing based on the language type (para 0130, processing acoustic cues at differing frequency/tones tied to the emotion); tuning the characteristics based on the desired results (preferences set by choosing an emotional category closest to the sample – see para 0121, and choosing/selecting the percentage that is ‘close enough’ – end of para 0121); speech characteristics being any one of speech, pitch, spacing, volume, etc. tied to the emotions and verbal expressions – para 0030, disclosing fundamental frequency, pitch, intensity, loudness, speaker rate, etc. etc. Therefore, it would have been obvious to one of ordinary skill in the art of emotion detection/extraction, from speech information, to enhance the system of Wang (GB 2571853) with the further processing of speech characteristics as taught by Patel et al (20120089396) because it would advantageously improve upon the end-user understanding/ perceptions (see Patel et al (20120089396), para 0102 – see perceptual improvements in the listed categories).
As per claim 19, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, wherein the one or more of the preferred speech characteristics specify a preferred tone, volume, pacing or pitch of synthesized speech outputs (see Wang (GB 2571853) , pp 8, top, inferring intonation/speed of speech; and Patel et al (20120089396) -- , para 0030, disclosing fundamental frequency, pitch, intensity, loudness, speaker rate, etc.).
As per claim 20, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, wherein one or more of the derived speech characteristics specify a discerned emotion or intent (see Wang (GB 2571853), deriving emotion and expressions from the input speech – middle pp 6, “extraction module of speech’s emotional characteristic parameters intended to extract the parameters with emotional characteristics in pre-process speech signal…” ).
As per claims 21-23,25, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the first device comprises a gaming console; transmitting the data packets over the network; the audio of the utterance of the first user is recorded via a player chat application that is implemented by the first device (see Wang (GB 2571853), pp6, wherein the environment is a mobile voice terminal tied to a virtual environment terminal (e.g, gaming console – claim 21) tied to a network for a 3D virtual environment for multiple users – abstract (claim 22,23); also utilizing video footage – pp4, second full paragraph – “video device” recording and displaying – see following paragraph “Specific Operations:…”).
As per claim 24, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, wherein the one or more derived speech characteristics comprise a discerned background noise that is associated with the audio (as, monitoring and processing, for noise in the signal – see Patel et al (20120089396), para 0062 – 0063).
As per claim 26, the combination of Wang (GB 2571853) in view of Patel et al (20120089396) teaches the method of claim 1, comprising receiving, by the first device, data indicating one or more values specified by one or more respective tunable knobs that were adjusted by the first user at the first device for predefining the one or more respective preferred speech characteristics for synthesizing speech outputs (see Wang (GB 2571853), allowing the operator to make adjustments on text/voice/video – pp 14, bottom “Upon selection…”, reflecting back to “Upon recording…text,voice, video; in view of the mapping to claim 1 above, the adjustments would follow through to the speech characteristics, in the combination of Wang (GB 2571853) in view of Patel et al (20120089396)).
Claims 27-35 are system claims that perform the commonly shared steps of method claims 1, 19-26 above and as such, claims 27-35 are similar in scope and content to claims 1,19-26; therefore, claims 27-35 are rejected under similar rationale as presented against claims 1,19-26 above. Furthermore, Wang (GB 2571853) teaches processors/storage devices, for executable step retrieval and processing (see Wang (GB 2571853), pp 6, middle – processor/memory.
Claims 36, 37 are non-transitory computer readable media claims whose instructions/steps are executed by a processor with said steps, are found throughout in method claims 1,19-26 above and as such, claims 36,37 are similar in scope and content to claims 1,19-26; therefore, claims 36,37 are rejected under similar rationale as presented against claims 1,19-26 above. Furthermore, Wang (GB 2571853) teaches processors/storage devices, for executable step retrieval and processing (see Wang (GB 2571853), pp 6, middle – processor/memory.
Response to Arguments
Applicant’s arguments with respect to claim(s) have been considered but are moot because the new ground of rejection refer to new citations/combinations not previously presented. Furthermore, examiner notes that applicants arguments are toward the amended claim language; examiner points to the further detailed explanations/mappings to the Wang (GB 2571853)/ Patel et al (20120089396) references.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Please note the reference cited on the PTO-892 form.
In further detail, examiner notes the following references pertaining to applicants spec/claim scope:
Mennicken et al (20210104220) teaches the selection of a persona from a plurality of personas based on the user -- Mennicken et al (20210104220) – see para 0082, wherein the user-selected characteristics to be used for specific users – end of para 0082.
Zeng et al (20210073526) teaches extraction of emotional information, from the information of the user (in fact, Zeng et al teaches extraction of emotional information from multiple modes – visual, audio, and text – see para 0053; and further, are compared for verification through a fusing process based on semantic meaning and maximized offset – para 0053, and in further detail para 0054, and the like).
Iwase et al (20200051545) teaches user selectable tones/types – para 0081
Yamagami et al (20090259475) teaches user selectable voice quality changes – para 0162
Sohn et al (20140022370) teaches storing of facial-emotion relationships for predicting from audio – para 0016
Socolof et al (20210097468) teaches analysis of sentiment clusters during a user interaction (para 0035, 0078).
Ahn et al (20140093849) teaches analysis and selection of estimated emotions from a collection of emotion vectors (para 0006).
Kang et al (20100121804) teaches estimating emotion using emotion vector parameters for comparison (para 0019).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Michael N Opsasnick/Primary Examiner, Art Unit 2658 07/19/2026