DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments with respect to Double Patenting rejection of claims 21-28, 30-38 and 40 have been considered and found persuasive, and the rejection has been withdrawn. Applicant filed a Terminal Disclaimer on 5/21/2026.
Applicant's arguments with respect to 35 U.S.C. 102 in regards to claims 21 and 31 have been considered but are moot due to new grounds of rejection necessitated by amendments. See detailed rejection below.
Applicant's arguments with respect to 35 U.S.C. 101 in regards to claims 21-28, 30-38 and 40 have been considered, however are not found to be persuasive due to the following reasons. Examiner respectfully disagrees with Applicant’s arguments because although claims 21 and 31 recite computer components such as a first device, a GUI, a device profile, and speech synthesis processing, merely performing an abstract idea on a computer does not make the claims patent eligible. The focus of the claims remain collecting user preference information, storing that information, retrieving it later, and applying it when generating synthesized speech. These are data management functions that can be characterized as organizing and using information, rather than an improvement to computer technology itself. The recited GUI, device profile and speech synthesis components are described only a functional level and perform their ordinary, well-understood functions. The claims do not recite a specific improvement to how speech synthesis is performed, how a device profile is structured, how speech parameters are determined, or how the GUI operates. Instead, these generic computer elements merely implement the abstract idea in a conventional technological environment and therefore do not integrate the judicial exception into a practical application.
Applicant’s reliance on the Specification and the decisions in Ex parte Desjardins and Ex parte Carmody is not persuasive. Unlike those cases, which involved claims directed to specific improvements in the operation or training of machine learning models, claims 21 and 31 do not improve the operation of a speech synthesis engine, a GUI, or any other computer technology. The alleged improvement identified by Applicant’s is an improved user experience by allowing speech preferences to be stored and reused. Improving user convenience or personalization, however, is not the same as improving the functioning of a computer or another technology. The claims simply use conventional computer components to automate the storage and later application of user selected speech characteristics, without reciting any particular technical solution or unconventional implementation. Accordingly, when considered as a whole, the claims remain directed to an abstract idea implemented using generic computer technology and do not include significantly more than the judicial exception under Step 2A or Step 2B of the Alice/Mayo framework.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21-42 are rejected under 35 U.S.C. 101.
Claims 21 and 31 are directed to an abstract idea to the basic abstract concept of storing user preferences and applying them to a later task. When we strip away the technical jargon, the claim is simply describing the process of asking a user how they want a computer voice to sound, remembering that choice for their specific device, and then using that saved choice the next time the device needs to speak. These are fundamental steps of collecting data, recognizing a user or device, and storing or retrieving information are considered abstract ideas. They represent routine organizational and data-gathering concepts that humans have performed mentally or manually for a very long time.
Simply taking an abstract idea and telling a computer to do it does not make it a patentable invention. For an abstract idea to be patentable, it must be integrated into a practical application that actually improves how the technology works. These claim does not describe a new, innovative way to generate synthesized speech or a technically improved computer interface. Instead, they relies on completely standard, off-the-shelf computer functions to do the job. It uses a generic GUI to receive the choice, basic computer memory to store the association, and standard "speech synthesis processing" to generate the voice.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device.
Dependent claims 22-30 and 32-42 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known in speech synthesis processing.
Claims 22 and 32, abstract concept of organizing information.
Claims 23 and 33, abstract mental process of sorting and categorizing data.
Claims 24 and 34, abstract idea of sending and updating data over a generic network.
Claims 25 and 35, the abstract idea of translating data from one format to another.
Claims 26 and 36, abstract mathematical algorithm.
Claims 27 and 37, the abstract step of gathering data.
Claims 28 and 38, a routine data-gathering step.
Claims 29 and 39, merely changes the informational content of the data being processed, which does not make the underlying idea less abstract.
Claims 30 and 40, does not transform the abstract idea into a patentable invention.
Claims 41 and 42, implement the same abstract idea of storing and applying user selected speech preferences.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 21-23, 25, 27-33, 35 and 37-40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Buntschuh (EP 0 762 384) in view of Almudafar-Depeyrot et al. (US 2018/0182373).
Claims 21 and 31,
Buntschuh teaches a computer-implemented method, comprising (“As shown in Fig. 1, there is illustrated an exemplary embodiment of a computer-based speech synthesizer system 02 that comprises a processing unit 07, a display screen terminal 08, input devices, e.g., a keyboard 10 and a mouse 12”; Buntschuh further states: “The processing unit 07 includes a processor 04 and a memory 06” [pg. 3] [Fig. 1]):
presenting, on a first device, a graphical user interface comprising a first element corresponding to a first characteristic of speech of a first synthesized voice and a second element corresponding to a second characteristic of speech of the first synthesized voice (“An exemplary embodiment of the present invention is written using the Tk toolkit to generate the present invention graphical user interface display 20 shown in Fig. 4 through X Windows”; The interface “comprises the following controls: parameter scales 22, a scrollable list 24 of named voices, a male voice button 26a, a female voice button 26b, an input box 28 … and a “say it” button 30”; The scales modify “pitchT, pitchR, pitchB, rate, front head, back head and aspiration”; each parameter scale 22 has a slider 22a…” [pg. 4] [Figs. 3-5]);
receiving, by the first device, a first user input corresponding to the first characteristic of speech (“The parameters scales are responsive to input from a user for altering current speech parameter values”; the user changes the value by “clicking on or selecting the “-“ or “+” buttons 22b and 22c, dragging the slider 22a, or clicking in the scale 22”; “any of these actions will trigger the occurrence of an event of interest to the present invention GUI” [pg. 2] [pg. 4] [Fig. 5]);
determining, using the first user input, first data representing the first characteristic (“the parameter scales 22 display a scale value 22d that corresponds to the relative position of the slider 22a within the range of the corresponding parameter scale 22”; “each time the sliders 22a are repositioned, the scale widget evaluates a Tcl command that causes the current speech parameter values to be updated with the scale value 22d”; Buntschuh also states that “it becomes necessary to update the current speech parameter value with the current scale value 22d … this is done by step 5d” [pgs. 4-5] [Fig. 5] [steps 5a-5d]);
after storing the association, receiving, by the first device, a second user input (“the present invention allows users to save current speech parameter values as a newly created named voice”; “this new named voice will be subsequently added to the list 24 and stored in the default initialization file with its associated speech parameters values”; thereafter, “the input box 28 is created with the Tk entry widget to permit the user to enter the test utterances, i.e., the text the user desires to have Bell Labs text-to-speech synthesizer convert to speech”; the user then presses return or clicks the “say it” button 30” [pg. 5] [Figs. 3, 4 and 8]);
determining output data responsive to the second user input (“this Tcl script forms and transmits to the Bell Labs text-to-speech synthesizer … the text string comprised of the series of escape sequences followed by the test utterances from the input box 28”; “the Tcl script first pairs the escape codes with their associated speech parameter values and then strings them together to form the series of escape sequence”; Buntschuh also states: “the formation means are operative to create a text string which includes the current speech parameter values”; “the text string may also include test utterances and escape codes” [pg. 5] [Figs. 2-3] [step 3c]); and
performing speech synthesis processing using the first data and the output data to determine synthesized speech data responsive to the second user input and corresponding to the first synthesized voice having the first characteristic (Buntschuh sends “a text string comprised of a series of escape sequences and text utterances to the Bell Labs text-to-speech synthesizer”; “the escape sequence are ASCII codes comprised of pairs of escape codes and associated speech parameter values”; the values identify “which speech parameters are to be set and the values to be assigned to each of the speech parameters”; “the test utterances represent the text to be converted to speech”; “upon receipt of the text string, the Bell Labs text-to-speech synthesizer will convert the test utterances to speech using a base synthesized voice altered according to the escape sequences” [pgs. 4-5] [Figs. 2-3]).
The difference between the prior art and the claimed invention is that Buntschuh does not explicitly teach storing an association between the first data and a device profile corresponding to the first device; after receiving the second user input, using the association to determine that the first data is to be used to respond to the second user input.
Almudafar teaches storing an association between the first data and a device profile corresponding to the first device (“some embodiments store TTS voice parameters within user profiles, such as TTS voice parameters described above with respect to Fig. 3”; “in various embodiments, system designers or system users are able to configure the TTS voice parameters in the user profile”; Almudafar further states: “In some embodiments, a system supports different types of voice-enabled devices, and each device has a device profile”; “some device have female-gendered TTS voice parameters and other devices have male-gendered TTS voice parameters” [0036-0037] [Figs. 4 & 13]);
after receiving the second user input, using the association to determine that the first data is to be used to respond to the second user input (“a parametric speech synthesis module 201 consumes input text and according to a set of TTS voice parameters, produces speech audio for a listener 102”; “a function module 406 produces the TTS voice parameters” by “transforming user profile attributes from a user profile 405 according to a model 407”; Almudafar further describes: “First, the system reads user profile attribute values 1301… Next, it executes a function on the user profile and model to produce TTS voice parameters 1303. Finally, the system performs parametric speech synthesis on text according to the TTS voice parameters 1304” [0036] [Figs. 4 and 13] [steps 1301-1304]).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Buntschuh with teachings of Almudafar by modifying the method and apparatus for modifying voice characteristics of synthesized speech as taught by Buntschuh to include storing an association between the first data and a device profile corresponding to the first device; after receiving the second user input, using the association to determine that the first data is to be used to respond to the second user input as taught by Almudafar for the benefit of providing an improved approach for adapting a voice to a particular customer for improving effectiveness of advertising (Almudafar [0004]).
Claims 22 and 32,
Buntschuh further teaches the computer-implemented method of claim 21, further comprising: determining the first characteristic corresponds to a first application with respect to the first device ([col. 8 line 57 to col. 9 line 1] once some voices have been created and stored, they can be used to process dialogue scripts or other applications);
associating the first data with the first application ([Fig. 9] [col. 9 lines 3-6] the preprocessor accesses data in step 9a from a voice file, as shown in Fig. 9, which contains a list of named voices and their associated speech parameter values); and
determining the second user input corresponds to the first application ([col. 9 lines 6-10] in steps 9b and 9c, the preprocessor filters out the bracket-enclosed speaker names and then replaces them with escape sequences formed using the speech parameter values associated with the named voices matching the speaker names),
wherein performing the speech synthesis processing using the first data is based at least in part on the second user input and the first data corresponding to the first application ([col. 9 lines 10-14] the escape sequences and the utterances are output in step 9d to the Bell Labs text-to-speech synthesizer to be converted to speech; the result is a spoken colloquy with different voices).
Claims 23 and 33,
Buntschuh further teaches the computer-implemented method of claim 21, further comprising: determining the first characteristic corresponds to a first application with respect to the first device ([col. 8 line 57 to col. 9 line 1] once some voices have been created and stored, they can be used to process dialogue scripts or other applications);
associating the first data with the first application (([Fig. 9] [col. 9 lines 3-6] the preprocessor accesses data in step 9a from a voice file, as shown in Fig. 9, which contains a list of named voices and their associated speech parameter values);
receiving a third user input corresponding to the second characteristic of speech to be associated with a second application with respect to the first device ([col. 1 lines 51-55] [col. 8 line 57 to col. 9 line 1] a virtual continuum of new voices created via user input (parameter scales) and storing them for use in various other applications);
determining, using the third user input, second data representing the second characteristic ([col. 1 lines 1-end] by manipulating the above mentioned speech parameters, a virtual continuum of new voice can be created via sliders to determine the current speech parameter values); and
associating the second data with the second application and with the first device ([col. 8 line 54 to col. 9 line 17] created voices are stored in a library (on the device memory) to be matched and used in multiple dialogue scripts or other applications).
Claims 25 and 35,
Buntschuh further teaches the computer-implemented method of claim 21, wherein determining the first data comprises: determining, using the first user input, a first value representing the first characteristic ([col. 6 lines 32-end] a scale value 22d that corresponds to the relative position of the slider 22a within the range of the corresponding parameter scale 22); and
determining, using the first value, encoded data representing the first characteristic, wherein the first data comprises the encoded data ([col. 5 lines 43-45] the escape sequences are ASCII codes comprised of pairs of escape codes and associated speech parameter values).
Claims 27 and 37,
Buntschuh further teaches the computer-implemented method of claim 21, wherein: the first user input was received by the first device in response to a touch interaction corresponding to the first element ([col. 6 lines 32-end] mouse click to drag the sliders via user interaction).
Claims 28 and 38,
Buntschuh further teaches the computer-implemented method of claim 27, further comprising: determining the first user input corresponds to manipulation of the first element from a first position to a second position ([col. 6 lines 32-end] use the mouse to drag the scale value from first position to second position (repositioning the slider)),
wherein determining the first data is based at least in part on the manipulation ([col. 6 lines 32-end] repositioning the slider).
Claims 29 and 39,
Buntschuh further teaches the computer-implemented method of claim 21, wherein the first characteristic of speech corresponds to a speech rate ([col. 6 lines 35-36] speech parameters includes speech rate).
Claims 30 and 40,
Buntschuh further teaches the computer-implemented method of claim 21, wherein the first characteristic of speech corresponds to an emotion ([col. 6 lines 35-36] speech parameters includes pitchT, pitchR, pitchB and aspiration).
Claim(s) 24 and 34 is/are rejected under 35 U.S.C. 103 as being unpatentable over Buntschuh (EP 0 762 384) in view of Almudafar-Depeyrot et al. (US 2018/0182373) and further in view of Kazan et al. (US 2011/0179149).
Claim 24 and 34,
Buntschuh further teaches the computer-implemented method of claim 21, further comprising, ([Summary of the invention] the following manipulable parameter scales: three pitches, front and rear head of the vocal tract, rate and aspiration (three or more different characteristics)).
The difference between the prior art and the claimed invention is that Buntschuh does not explicitly teach after performing the speech synthesis processing: receiving, from a second device; determining the second device is associated with the first device; determining modified first data representing the third characteristic; and storing second data associating the modified first data with the first device.
Kazan teaches after performing the speech synthesis processing: receiving, from a second device ([0004] a second change to application settings on a second device of the one or more additional computing devices);
determining the second device is associated with the first device ([0036-0037] a record of the computing devices across which application settings are roamed can be maintained at computing device 200 and devices identified via computing device from the user logs into a remote service);
determining modified first data representing the third characteristic ([Abstract] [0070] application setting changes received from other computing devices across which application settings are roamed, and are incorporated into the application settings of the computing device as discussed above); and
storing second data associating the modified first data with the first device ([0016] whenever a change to a roamed application setting is made on one of computing devices 102, 104, and 106, it is automatically communicated to and saved by the other computing devices 102, 104, and 106).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Buntschuh with teachings of Kazan by modifying the method and apparatus for modifying voice characteristics of synthesized speech as taught by Buntschuh to after performing the speech synthesis processing: receiving, from a second device; determining the second device is associated with the first device; determining modified first data representing the third characteristic; and storing second data associating the modified first data with the first device as taught by Kazan include for the benefit of performing the same application customization on each of the multiple devices ([0002] Kazan).
Claim(s) 26 and 36 is/are rejected under 35 U.S.C. 103 as being unpatentable over Buntschuh (EP 0 762 384) in view of Almudafar-Depeyrot et al. (US 2018/0182373) and further in view of Reber et al. (US 2019/0304435).
Claims 26 and 36,
Buntschuh teaches all the limitations in claim 25. The difference between the prior art and the claimed invention is that Buntschuh does not explicitly teach wherein performing the speech synthesis processing comprises processing the encoded data using a neural network speech synthesis processing component to determine the synthesized speech data.
Reber teaches wherein performing the speech synthesis processing comprises processing the encoded data using a neural network speech synthesis processing component to determine the synthesized speech data ([Abstract] to provide analysis and conversion of text into input vectors, each having at least a base frequency, f.sub.0, a phenome duration, and a phoneme sequence that is processed by a signal generation unit of the back-end subsystem; the signal generation unit includes the neural network interacting with a pre-existing knowledgebase of phenomes to generate audible speech from the input vectors).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Buntschuh with teachings of Reber by modifying the method and apparatus for modifying voice characteristics of synthesized speech as taught by Buntschuh to include wherein performing the speech synthesis processing comprises processing the encoded data using a neural network speech synthesis processing component to determine the synthesized speech data as taught by Reber for the benefit of improving the quality of the generated audible speech signals ([Abstract] Reber).
Claim(s) 41-42 is/are rejected under 35 U.S.C. 103 as being unpatentable over Buntschuh (EP 0 762 384) in view of Almudafar-Depeyrot et al. (US 2018/0182373) and further in view of Qian et al. (US 2010/0066742).
Claims 41 and 42,
Buntschuh further teaches a third set of elements of the graphical user interface, the third set of elements corresponding to vocal characteristics (Buntschuh teaches that its GUI includes parameter scales for “three pitches, front and rear head of the vocal tract, rate and aspiration” [pgs. 2-5] [Figs. 3-5]).
The difference between the prior art and the claimed invention is that Buntschuh does not explicitly teach wherein the graphical user interface further comprises: a first set of elements of the graphical user interface, the first set of elements corresponding to speech styles; a second set of elements of the graphical user interface, the second set of elements corresponding to speech labels; a fourth set of elements of the graphical user interface, the fourth set of elements corresponding to phoneme characteristics; and a fifth set of elements of the graphical user interface, the fifth set of elements corresponding to feedback control elements.
Qian teaches wherein the graphical user interface further comprises: a first set of elements of the graphical user interface, the first set of elements corresponding to speech styles (Qian teaches a “visual interface” for supervising a speech-synthesis system to generate user-selected prosody, including “emotions, intonations and speaking styles” [0016-0018] [0022-0023] [Figs. 1 and 3]);
a second set of elements of the graphical user interface, the second set of elements corresponding to speech labels (Qian teaches that input text is converted into “a sequence of contextual labels through a text analysis component 240” [0027] [0030-0031] [Figs. 2 and 3]);
a fourth set of elements of the graphical user interface, the fourth set of elements corresponding to phoneme characteristics (Qian teaches that the interface permits changing prosody for portions of speech that “may comprise a phoneme… and/or a sentence” [0008] [0017] [0030-0032] [Figs. 3-4]); and
a fifth set of elements of the graphical user interface, the fifth set of elements corresponding to feedback control elements (Qian teaches that its interface “controls output to a speaker 112” for “replaying the initial speech and/or the modified prosody speech” [0020-0021] [0040-0047] [Figs. 1, 3 and 4 steps 408-421]).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Buntschuh with teachings of Qian by modifying the method and apparatus for modifying voice characteristics of synthesized speech as taught by Buntschuh to include wherein the graphical user interface further comprises: a first set of elements of the graphical user interface, the first set of elements corresponding to speech styles; a second set of elements of the graphical user interface, the second set of elements corresponding to speech labels; a fourth set of elements of the graphical user interface, the fourth set of elements corresponding to phoneme characteristics; and a fifth set of elements of the graphical user interface, the fifth set of elements corresponding to feedback control elements as taught by Qian for the benefit of ([Abstract] Reber).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
SHREYANS A. PATEL
Primary Examiner
Art Unit 2653
/SHREYANS A PATEL/ Examiner, Art Unit 2659