DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments/ Remarks
Applicant’s arguments filed 6/18/2026 with respect to claim(s) 1 to 20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Specification
Objections to the specification have been addressed and corrected. Objection is withdrawn.
Claim Objections
Objections to the claims have been addressed and corrected. Objection is withdrawn.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 2 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 describes, “identify a known contact assigned to the second caller, wherein the known contact is assigned relative to the first caller”. However, claim 1 describes, “identify an unknown contact assigned to the second caller, wherein the contact is assigned relative to the first caller”. Since claim 1 describes that the second caller is an unknown contact, relative to the first caller, it is unclear as to how the second caller becomes a known contact, relative to the first caller, in claim 2.
For the purpose of this examination the “second caller” will be interpreted as a “third caller” in claim 2.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 2, 4, 5, 8, 10, 11, 12, 14, 15, 16, 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Efrati, Tzahi et al. (US-20160210959-A1) hereinafter Efrati in view of HENROTTE; Christophe et al. (US 20230410825 A1) hereinafter Henrotte.
Regarding claim 1, Efrati teaches:
(Currently Amended) A system comprising: a processor configured to: identify a first type call between a first caller and a second caller;
“In some embodiments, the participant metadata 301 may include information stored with respect to a user profile associated with a unique identifier (e.g., caller ID, messaging system identifier, email address, account identifier, IP address, name, and the like) of the user. The user profile may be maintained in a database by the service provider. For example, according to exemplary embodiments, the participant metadata may include such information as the telephone number of the participant, a caller ID of the participant, the city, state and country of the participant, a voice sample, social media information, contacts and the like. In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include previous dialect information detected or used on a previous history of call(s). ...” (Efrati [0045])
identify an unknown contact assigned to the second caller, wherein the contact is assigned relative to the first caller;
“ According to one embodiment, once the call has begun and participants (for example, participants 210-216) of the conversation have started speaking, the dialect modification apparatus 200 receives a source audio signal associated with each participant. The detection module 202 then detects the dialect of each participant based on one or more of their voices in the call, their caller ID, their user ID if using a VoIP network, and associated metadata. In other words, the detection module 202 detects a source dialect for each participant. In some instances, the metadata may further contain location information for the origination of the call, social media profile information, contact information, destination of the call or the like. Taken together, this data can strongly predict a particular user's dialect. ” (Efrati [0034]).
“ In some instances, a caller identification number (CLID) can be used to retrieve a caller's location and as a result, the dominant dialect in that area. In addition, if the service provider 201 stores a user address book, then the CLID can be found in a contact which, in turn, may provide more information such as physical address or location of the caller. ” (Efrati [0035]).
Wherein the system gathers information about the person calling based on some parameters like location associated to their caller ID from the caller, indicating that the caller is not known to the system and therefore to the user.
retrieve a seed associated with the unknown contact;
“ According to one embodiment, once the call has begun and participants (for example, participants 210-216) of the conversation have started speaking, the dialect modification apparatus 200 receives a source audio signal associated with each participant. The detection module 202 then detects the dialect of each participant based on one or more of their voices in the call, their caller ID, their user ID if using a VoIP network, and associated metadata. In other words, the detection module 202 detects a source dialect for each participant. …” (Efrati [0034]).
“ … In some embodiments, the call may be routed to any customer service or sales agent, for example, but the agent's voice would be modulated accordingly. In some embodiments, the agent's voice would be modulated to a target dialect based on a preference of the caller, or based on a customer satisfaction ratings or sales revenue generation data associated with particular dialects. ” (Efrati [0053]).
Wherein the voice seed is selected for each participant utilizing metadata collected from the call, and by form of example Efrati shows a customer service representative that change their voice seed to a particular caller, wherein the caller receive a modified version of the voice of the customer support representative.
And generate a synthetic voice based in part on the voice seed and an input from the first caller.
“According to another embodiment, the conversion module 204 passes the voice call content 400 directly to the modulation module 206 without performing speech-to-text conversion. The modulation module 206 then parses the voice call content 400 into various phonemes. The difference between the phonemes of the participant and the phonemes of the target dialect are determined, and based on the speech profile 300, the modulation module 206 modulates portions of the voice call content 400 to generate a modulated voice 404. The modulated voice 404 is modulated according to the predetermined or chosen one or more target dialects. The modulated voice 404 is then relayed to the appropriate call participants, based on which dialect the recipients are programmed to hear.” (Efrati [0052]).
Efrati does not disclose explicitly, but Henrotte teaches:
Masking (voice seed)
“ A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. ... ” (Henrotte [Abstract]).
“ By virtue of this method, the voice is masked, making it possible to respond to the requirement to protect the one or more speakers, since the method is easily able to be implemented in any first equipment involved in the audio acquisition and processing chain. At the same time, the method makes it possible to have a final rendering that remains intelligible, that is to say that is neither a “Mickey Mouse”™ voice nor a “Darth Vader” ™ voice, due to the two alterations applied to each audio segment, which produce modifications in the frequency content that are in a direction contrary to one another. Indeed, a rising effect (towards high-pitched tones) is applied to one of the two alterations and a falling effect (towards low-pitched tones) is applied to the other of the two alterations, such that these two effects combine from the point of view of the frequency content of the audio segment under consideration. The resulting masked audio segment possesses frequency content that remains overall closer, over the spectral dynamic range, to that of the original audio segment, despite the voice masking that is obtained. ” (Henrotte [0023]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Efrati the capability to use the voice changing features as a masking for the voice of the user. The benefit and motivation of such modification is discussed by Henrotte in the following portion: ”The method makes it possible to mask the voice of a speaker for the purpose of protecting their identity and/or their privacy.” (Henrotte [0051]).
Regarding claim 2, the rejection of claim 1 is incorporated, furthermore Efrati teaches:
(Currently Amended) The system of claim 1, wherein the processor is further configured to: identify a second type call between the first caller and a third caller;
“In some embodiments, the participant metadata 301 may include information stored with respect to a user profile associated with a unique identifier (e.g., caller ID, messaging system identifier, email address, account identifier, IP address, name, and the like) of the user. The user profile may be maintained in a database by the service provider. For example, according to exemplary embodiments, the participant metadata may include such information as the telephone number of the participant, a caller ID of the participant, the city, state and country of the participant, a voice sample, social media information, contacts and the like. In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include previous dialect information detected or used on a previous history of call(s). ...” (Efrati [0045])
identify a known contact assigned to the second caller, wherein the known contact is assigned relative to the first caller;
Note: as stated above, for purpose of this examination the “second caller” of this claim would be considered as the “third caller”.
“In some embodiments, the participant metadata 301 may include information stored with respect to a user profile associated with a unique identifier (e.g., caller ID, messaging system identifier, email address, account identifier, IP address, name, and the like) of the user. The user profile may be maintained in a database by the service provider. For example, according to exemplary embodiments, the participant metadata may include such information as the telephone number of the participant, a caller ID of the participant, the city, state and country of the participant, a voice sample, social media information, contacts and the like. In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include previous dialect information detected or used on a previous history of call(s). ...” (Efrati [0045])
Wherein the user during a phone call receiving the name of the caller is related to a caller that is in their contact, therefore could be considered known to the user. Contrary to that, the user would receive a caller ID of the caller.
retrieve a configured voice seed associated with the known contact;
“FIG. 3 is a block diagram detailing the operation of the detection module 202 in accordance with exemplary embodiments of the present invention. As described above, the detection module 202 parses the voice content of all participants and retrieves an associate's speech profile associated with each participant. According to FIG. 3, the detection module 202 retrieves a speech profile 300 from the datastore 208 based on the detected user dialect. In some embodiments, the selection of a speech profile is greatly enhanced by participant metadata 301 to increase the accuracy of the selected speech profile.” (Efrati [0044]).
and generate a configured synthetic voice based in part on the configured voice seed and an input from the first caller.
“According to another embodiment, the conversion module 204 passes the voice call content 400 directly to the modulation module 206 without performing speech-to-text conversion. The modulation module 206 then parses the voice call content 400 into various phonemes. The difference between the phonemes of the participant and the phonemes of the target dialect are determined, and based on the speech profile 300, the modulation module 206 modulates portions of the voice call content 400 to generate a modulated voice 404. The modulated voice 404 is modulated according to the predetermined or chosen one or more target dialects. The modulated voice 404 is then relayed to the appropriate call participants, based on which dialect the recipients are programmed to hear.” (Efrati [0052]).
Regarding claim 4, the rejection of claim 1 is incorporated, furthermore Efrati teaches:
(Currently Amended) The system of claim 1, wherein the processor is further configured to modify the seed based on a modification request received from a device associated with the first caller.
“…In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include a user selection of a first dialect they would like to voice to be modulated to (i.e., how a user would like to sound to other participants), and/or a user selection of a second dialect that they would like the other participant(s) voice to be modulated to (how a user would like other participants to sound to them)...” (Efrati [0045]).
Efrati does not disclose explicitly, but Henrotte teaches:
Masking (voice seed)
“ A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. ... ” (Henrotte [Abstract]).
“ By virtue of this method, the voice is masked, making it possible to respond to the requirement to protect the one or more speakers, since the method is easily able to be implemented in any first equipment involved in the audio acquisition and processing chain. At the same time, the method makes it possible to have a final rendering that remains intelligible, that is to say that is neither a “Mickey Mouse”™ voice nor a “Darth Vader” ™ voice, due to the two alterations applied to each audio segment, which produce modifications in the frequency content that are in a direction contrary to one another. Indeed, a rising effect (towards high-pitched tones) is applied to one of the two alterations and a falling effect (towards low-pitched tones) is applied to the other of the two alterations, such that these two effects combine from the point of view of the frequency content of the audio segment under consideration. The resulting masked audio segment possesses frequency content that remains overall closer, over the spectral dynamic range, to that of the original audio segment, despite the voice masking that is obtained. ” (Henrotte [0023]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Efrati the capability to use the voice changing features as a masking for the voice of the user. The benefit and motivation of such modification is discussed by Henrotte in the following portion: ”The method makes it possible to mask the voice of a speaker for the purpose of protecting their identity and/or their privacy.” (Henrotte [0051]).
Regarding claim 5, the rejection of claim 4 is incorporated, furthermore Efrati teaches:
(Currently Amended) The system of claim 4, wherein modifying the voice seed further comprises: receiving one or more seed generation parameters from the device;
“…In some embodiments, each participant may require that incoming voice calls be modulated to the participant's dialect. For example, the first participant 210 may configure it so that their personal dialect is the target dialect for all incoming calls to the first participant 210. Similarly, on the same call, third participant 214 may configure it so that all incoming voice is modulated to the dialect of participant 214.” (Efrati [0040]).
“The speech profile 300 is comprised of regional information 302, dialect information 304, phonetic transformation information 306 and acoustic information 308…” (Efrati [0046]).
“At step 606, one or more target dialects are chosen for at least one of the one or more participants. Namely, it is not necessary that only one target dialect be selected for all participants. The dialect modification module 506 can be configured to enable each participant to hear other participants' speech in their dialect…” (Efrati [0063]).
And generating the voice seed from input of the seed generation parameters.
“…In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include a user selection of a first dialect they would like to voice to be modulated to (i.e., how a user would like to sound to other participants), and/or a user selection of a second dialect that they would like the other participant(s) voice to be modulated to (how a user would like other participants to sound to them)…” (Efrati [0045]).
Efrati does not disclose explicitly, but Henrotte teaches:
Masking (voice seed)
“ A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. ... ” (Henrotte [Abstract]).
“ By virtue of this method, the voice is masked, making it possible to respond to the requirement to protect the one or more speakers, since the method is easily able to be implemented in any first equipment involved in the audio acquisition and processing chain. At the same time, the method makes it possible to have a final rendering that remains intelligible, that is to say that is neither a “Mickey Mouse”™ voice nor a “Darth Vader” ™ voice, due to the two alterations applied to each audio segment, which produce modifications in the frequency content that are in a direction contrary to one another. Indeed, a rising effect (towards high-pitched tones) is applied to one of the two alterations and a falling effect (towards low-pitched tones) is applied to the other of the two alterations, such that these two effects combine from the point of view of the frequency content of the audio segment under consideration. The resulting masked audio segment possesses frequency content that remains overall closer, over the spectral dynamic range, to that of the original audio segment, despite the voice masking that is obtained. ” (Henrotte [0023]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Efrati the capability to use the voice changing features as a masking for the voice of the user. The benefit and motivation of such modification is discussed by Henrotte in the following portion: ”The method makes it possible to mask the voice of a speaker for the purpose of protecting their identity and/or their privacy.” (Henrotte [0051]).
Regarding claim 8, the rejection of claim 2 is incorporated, furthermore Efrati teaches:
(Currently Amended) The system of claim 2, wherein the
“In some embodiments, the participant metadata 301 may include information stored with respect to a user profile associated with a unique identifier (e.g., caller ID, messaging system identifier, email address, account identifier, IP address, name, and the like) of the user. The user profile may be maintained in a database by the service provider. For example, according to exemplary embodiments, the participant metadata may include such information as the telephone number of the participant, a caller ID of the participant, the city, state and country of the participant, a voice sample, social media information, contacts and the like. In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include previous dialect information detected or used on a previous history of call(s). In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include a user selection of a first dialect they would like to voice to be modulated to (i.e., how a user would like to sound to other participants), and/or a user selection of a second dialect that they would like the other participant(s) voice to be modulated to (how a user would like other participants to sound to them). In some embodiments, an identifier of one or more of the participants may be used to lookup participant metadata 301 in a user database to determine dialects used in previous calls by the user. Those of ordinary skill in the art will recognize that the participant metadata 301 may include any information available from a service provider which aids in identifying the user's region or dialect information as described above. ” (Efrati [0045]).
Efrati does not disclose explicitly, but Henrotte teaches:
Masking (synthetic voice, voice seed)
“ A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. ... ” (Henrotte [Abstract]).
“ By virtue of this method, the voice is masked, making it possible to respond to the requirement to protect the one or more speakers, since the method is easily able to be implemented in any first equipment involved in the audio acquisition and processing chain. At the same time, the method makes it possible to have a final rendering that remains intelligible, that is to say that is neither a “Mickey Mouse”™ voice nor a “Darth Vader” ™ voice, due to the two alterations applied to each audio segment, which produce modifications in the frequency content that are in a direction contrary to one another. Indeed, a rising effect (towards high-pitched tones) is applied to one of the two alterations and a falling effect (towards low-pitched tones) is applied to the other of the two alterations, such that these two effects combine from the point of view of the frequency content of the audio segment under consideration. The resulting masked audio segment possesses frequency content that remains overall closer, over the spectral dynamic range, to that of the original audio segment, despite the voice masking that is obtained. ” (Henrotte [0023]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Efrati the capability to use the voice changing features as a masking for the voice of the user. The benefit and motivation of such modification is discussed by Henrotte in the following portion: ”The method makes it possible to mask the voice of a speaker for the purpose of protecting their identity and/or their privacy.” (Henrotte [0051]).
Regarding claim 10, the rejection of claim 4 is incorporated, furthermore Efrati teaches:
(Currently Amended) The system of claim 4, wherein generating the synthetic voice based in part on the voice seed and an input from the first caller comprises retaining one or more voiceprint characteristics of the input from the first caller.
“In some embodiments, this phoneme conversion is performed in real-time, or with a slight delay based on processing time of the underlying hardware of the dialect modification apparatus 200, preserving the identity of each participant while making each participant easier to understand to others. In some embodiments, participants who would like to hear other participants without voice modulation may disable any voice modulation generated by the dialect modification apparatus 200.” (Efrati [0042]).
Efrati does not disclose explicitly, but Henrotte teaches:
Masking (synthetic voice, voice seed)
“ A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. ... ” (Henrotte [Abstract]).
“ By virtue of this method, the voice is masked, making it possible to respond to the requirement to protect the one or more speakers, since the method is easily able to be implemented in any first equipment involved in the audio acquisition and processing chain. At the same time, the method makes it possible to have a final rendering that remains intelligible, that is to say that is neither a “Mickey Mouse”™ voice nor a “Darth Vader” ™ voice, due to the two alterations applied to each audio segment, which produce modifications in the frequency content that are in a direction contrary to one another. Indeed, a rising effect (towards high-pitched tones) is applied to one of the two alterations and a falling effect (towards low-pitched tones) is applied to the other of the two alterations, such that these two effects combine from the point of view of the frequency content of the audio segment under consideration. The resulting masked audio segment possesses frequency content that remains overall closer, over the spectral dynamic range, to that of the original audio segment, despite the voice masking that is obtained. ” (Henrotte [0023]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Efrati the capability to use the voice changing features as a masking for the voice of the user. The benefit and motivation of such modification is discussed by Henrotte in the following portion: ”The method makes it possible to mask the voice of a speaker for the purpose of protecting their identity and/or their privacy.” (Henrotte [0051]).
Regarding claim 11, Efrati teaches:
(Currently Amended) A method including operations executed by one or more processors, the operations comprising:
identifying a first type call between a first caller and a second caller
identifying an unknown contact assigned to the second caller, wherein the contact is assigned relative to the associated with second caller;
“ According to one embodiment, once the call has begun and participants (for example, participants 210-216) of the conversation have started speaking, the dialect modification apparatus 200 receives a source audio signal associated with each participant. The detection module 202 then detects the dialect of each participant based on one or more of their voices in the call, their caller ID, their user ID if using a VoIP network, and associated metadata. In other words, the detection module 202 detects a source dialect for each participant. In some instances, the metadata may further contain location information for the origination of the call, social media profile information, contact information, destination of the call or the like. Taken together, this data can strongly predict a particular user's dialect. ” (Efrati [0034]).
“ In some instances, a caller identification number (CLID) can be used to retrieve a caller's location and as a result, the dominant dialect in that area. In addition, if the service provider 201 stores a user address book, then the CLID can be found in a contact which, in turn, may provide more information such as physical address or location of the caller. ” (Efrati [0035]).
Wherein the system gathers information about the person calling based on some parameters like location associated to their caller ID from the caller, implying that the caller is not known to the system and therefore to the user.
retrieving a voice seed associated with the contact;
“ FIG. 3 is a block diagram detailing the operation of the detection module 202 in accordance with exemplary embodiments of the present invention. As described above, the detection module 202 parses the voice content of all participants and retrieves an associate's speech profile associated with each participant. According to FIG. 3, the detection module 202 retrieves a speech profile 300 from the datastore 208 based on the detected user dialect. In some embodiments, the selection of a speech profile is greatly enhanced by participant metadata 301 to increase the accuracy of the selected speech profile. ” (Efrati [0044]).
And generating a synthetic voice based in part on applying the seed and an input from the first caller.
“ According to another embodiment, the conversion module 204 passes the voice call content 400 directly to the modulation module 206 without performing speech-to-text conversion. The modulation module 206 then parses the voice call content 400 into various phonemes. The difference between the phonemes of the participant and the phonemes of the target dialect are determined, and based on the speech profile 300, the modulation module 206 modulates portions of the voice call content 400 to generate a modulated voice 404. The modulated voice 404 is modulated according to the predetermined or chosen one or more target dialects. The modulated voice 404 is then relayed to the appropriate call participants, based on which dialect the recipients are programmed to hear. ” (Efrati [0052]).
Efrati does not disclose explicitly, but Henrotte teaches:
Masking (voice seed, synthetic voice)
“ A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. ... ” (Henrotte [Abstract]).
“ By virtue of this method, the voice is masked, making it possible to respond to the requirement to protect the one or more speakers, since the method is easily able to be implemented in any first equipment involved in the audio acquisition and processing chain. At the same time, the method makes it possible to have a final rendering that remains intelligible, that is to say that is neither a “Mickey Mouse”™ voice nor a “Darth Vader” ™ voice, due to the two alterations applied to each audio segment, which produce modifications in the frequency content that are in a direction contrary to one another. Indeed, a rising effect (towards high-pitched tones) is applied to one of the two alterations and a falling effect (towards low-pitched tones) is applied to the other of the two alterations, such that these two effects combine from the point of view of the frequency content of the audio segment under consideration. The resulting masked audio segment possesses frequency content that remains overall closer, over the spectral dynamic range, to that of the original audio segment, despite the voice masking that is obtained. ” (Henrotte [0023]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Efrati the capability to use the voice changing features as a masking for the voice of the user. The benefit and motivation of such modification is discussed by Henrotte in the following portion: ”The method makes it possible to mask the voice of a speaker for the purpose of protecting their identity and/or their privacy.” (Henrotte [0051]).
Regarding claim 12, the rejection of claim 11 is incorporated, furthermore arguments analogous to claim 2 are applicable.
Regarding claim 14, the rejection of claim 11 is incorporated, furthermore arguments analogous to claim 4 are applicable.
Regarding claim 15, the rejection of claim 14 is incorporated, furthermore arguments analogous to claim 5 are applicable.
Regarding claim 16, the rejection of claim 15 is incorporated, furthermore Efrati teaches:
(Original) The method of claim 15, wherein the plurality of seed generation parameters includes audio accessibility parameters.
“The method begins at step 602 and proceeds to step 604. At step 604, the detection module 508 detects the dialect of one or more voice call participants. The detected dialects may be all the same, or each one may differ. The goal is to align the dialects so that all participants may understand each other. The dialects are detected based on a received source audio signal associated with at least one participant. The source audio signal comprises a voice of the at least one participant.” (Efrati [0062]).
“At step 606, one or more target dialects are chosen for at least one of the one or more participants. Namely, it is not necessary that only one target dialect be selected for all participants. The dialect modification module 506 can be configured to enable each participant to hear other participants' speech in their dialect…” (Efrati [0063]).
Regarding claim 18, the rejection of claim 12 is incorporated, furthermore arguments analogous to claim 2 are applicable.
Regarding claim 20, arguments analogous to claim 1 are applicable, furthermore Efrati teaches:
A non-transitory computer-readable medium embodying program code that, when executed by one or more processors, causes the processors to perform operations
“ The memory 504, or computer readable medium, stores non-transient processor-executable instructions and/or data that may be executed by and/or used by the processor 502. These processor-executable instructions may comprise firmware, software, and the like, or some combination thereof. Modules having processor-executable instructions that are stored in the memory 504 comprise an dialect modification module 506 and a datastore 514. The dialect modification module 506 further comprises a detection module 508, a conversion module 510 and a modulation module 512. Speech profiles 515 and voice samples 516 of the various participants in a voice call may also be stored in memory 504. In other instances, the speech profiles and voice samples are stored in a cloud storage for access and retrieval.” (Efrati [0055]).
Claim(s) 3, 7, 13 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Efrati in view of Henrotte, in further view of Guedalia, Jacob et al. (US-8756328-B2) hereinafter Guedalia.
Regarding claim 3, the rejection of claim 2 is incorporated, furthermore Efrati teaches:
receiving, from the device, a request to generate an associated voice seed;
“In some embodiments, the participant metadata 301 may include information stored with respect to a user profile associated with a unique identifier (e.g., caller ID, messaging system identifier, email address, account identifier, IP address, name, and the like) of the user. The user profile may be maintained in a database by the service provider. For example, according to exemplary embodiments, the participant metadata may include such information as the telephone number of the participant, a caller ID of the participant, the city, state and country of the participant, a voice sample, social media information, contacts and the like. In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include previous dialect information detected or used on a previous history of call(s). In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include a user selection of a first dialect they would like to voice to be modulated to (i.e., how a user would like to sound to other participants), and/or a user selection of a second dialect that they would like the other participant(s) voice to be modulated to (how a user would like other participants to sound to them). In some embodiments, an identifier of one or more of the participants may be used to lookup participant metadata 301 in a user database to determine dialects used in previous calls by the user. Those of ordinary skill in the art will recognize that the participant metadata 301 may include any information available from a service provider which aids in identifying the user's region or dialect information as described above. ” (Efrati [0045]).
generating a voice seed associated with the selected contact;
“In some embodiments, the participant metadata 301 may include information stored with respect to a user profile associated with a unique identifier (e.g., caller ID, messaging system identifier, email address, account identifier, IP address, name, and the like) of the user. The user profile may be maintained in a database by the service provider. For example, according to exemplary embodiments, the participant metadata may include such information as the telephone number of the participant, a caller ID of the participant, the city, state and country of the participant, a voice sample, social media information, contacts and the like. In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include previous dialect information detected or used on a previous history of call(s). In some embodiments, the participant metadata 301 retrieved from a user profile stored in a database maintained by the service provider may include a user selection of a first dialect they would like to voice to be modulated to (i.e., how a user would like to sound to other participants), and/or a user selection of a second dialect that they would like the other participant(s) voice to be modulated to (how a user would like other participants to sound to them). In some embodiments, an identifier of one or more of the participants may be used to lookup participant metadata 301 in a user database to determine dialects used in previous calls by the user. Those of ordinary skill in the art will recognize that the participant metadata 301 may include any information available from a service provider which aids in identifying the user's region or dialect information as described above. ” (Efrati [0045]).
and storing the voice seed in a voice key database.
“ Alternatively, a sample of the participant's voice is taken from the source audio signal by the detection module 202 and compared to existing dialects stored in a datastore 208. The datastore 208 may be a relational database, or other type of data storing service. In some instances, the datastore 208 may be located locally or remotely from the dialect modification apparatus 200. The datastore 208 stores dialects samples, speech profiles and other data related to voice modification during a call. In some embodiments, a speech profile is returned to the detection module 202 based on the dialect matched from the datastore 208. ” (Efrati [0037]).
Efrati does not teach, but Guedalia teaches:
(Currently Amended) The system of claim 2, wherein, the configured voice seed is generated via operations comprising:
transmitting, to a device associated with the first caller, a user a contact list;
“Other embodiments are directed to an apparatus wherein the display comprises a screen that provides a visual display of a plurality of contacts from said contacts list and permits perception of a state of presence of said contact.” (Guedalia, column 7 lines 60 to 63).
receiving, from the device, a selected contact;
“Still other embodiments are directed to a signaling system for establishing communication between a first mobile telephony device coupled to a mobile telephony network and a second communication device coupled to a data network, including first communication means for signaling communication between said first mobile telephony device and a server; a data storage and retrieval means, coupled to said server, for storing and maintaining a server contacts list of a plurality of contacts associated with said first mobile telephony device; a mobile contacts list correlated with said server contacts list and indicative of a state of information in said server contacts list, said mobile contacts list being accessible by said first mobile telephony device to provide a selected one or more contacts from said mobile contacts list to said server; and a second communication means for signaling communication between said second communication device and said server according to an address correlation at said server correlating said selected one or more contacts received over said mobile telephony network with a corresponding data network address of said second communication device. ” (Guedalia, column 7 lines 24 to 43).
“In related embodiments, the above system comprises said first communication device being a mobile telephone comprising a processor and software that runs on the processor. ” (Guedalia, column 8 lines 39 to 41).
“Still other embodiments include said first communication device having a smart phone comprising a processor and software that runs on the processor. ” (Guedalia, column 8 lines 42 to 44).
transmitting, to the device, a voice seed generator interface;
“Other embodiments are directed to an apparatus wherein the selector comprises a hardware user interface element that is constructed to receive an input from a user to select said contact from said stored contacts list.” (Guedalia, column 7 lines 64 to 67).
It would have been obvious for someone of ordinary skill in the art before the effective filing
date of the claimed invention to have modified Efrati to incorporate the teachings of Guedalia to include a contact list with the capabilities of providing a contact from the list and a graphical user interface for the voice modifier. The motivation and some of the benefits to include the contact list is discussed by Guedalia in: ”One aspect of the present invention allows users to integrate multiple contact lists stored on different devices. Generally, contact lists associate contact information (contact name, alias, etc.) with the network address of the contact. For instance, a contact list stored on the cellular phone may associate a contact Joe Smith with the phone number 617-123-1234. Similarly a contact list stored in the VoIP device may associate a contact "Smith" with the Internet Protocol address "66.249.64.15."” (Guedalia, column 17 lines 52 to 60).
Regarding claim 7, the rejection of claim 3 is incorporated, Efrati does not teach:
(Currently Amended) The system of claim 3, wherein the processor is further configured to restrict access to the voice seed generator interface by authenticating a user's access request received from the device.
On the other hand, Guedalia teaches:
“…This may involve an authentication sequence whereby device D1 and/or user U1 provide a user name or a password to server SVR. Also, the identity of device D1 may be transmitted through a serial number or other coded hardware and/or software scheme that identifies the processor, a key, or software or other token on device D1. Server SVR may look up the authentication log on information from device D1/user U1 directly, e.g. on a lookup table, or using an authentication server or client software on or coupled to or accessible to server SVR.” (Guedalia, column 11 lines 32 to 41).
It would have been obvious for someone of ordinary skill in the art before the effective filing
date of the claimed invention to have modified Efrati to incorporate the teachings of Guedalia to include instructions for the processor to require user authentication on the device to access the interface. The motivation and some of the benefits to include an authentication process is discussed by Guedalia in: “Device D1 and/or user U1 then "logs on" to server SVR over the portions of the communication path or network between device D1 and server SVR. This process is generally known to those skilled in the art and involves any of a number of authentication steps so that server SVR can determine the identity of device D1 and/or its user U1 to an acceptable degree of certainty.” (Guedalia, column 11 lines 26 to 32).
Regarding claim 13, the rejection of claim 12 is incorporated, furthermore arguments analogous to claim 3 are applicable.
Regarding claim 17, the rejection of claim 13 is incorporated, furthermore arguments analogous to claim 7 are applicable.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Efrati in view of Henrotte, in further view of Devitt; Jason A. et al. (US 8626137 B1) hereinafter Devitt
Regarding claim 6, the rejection of claim 1 is incorporated, furthermore Efrati in view of Henrotte does not teach but Devitt teaches:
(Currently Amended) The system of claim 1, wherein identifying an unknown contact assigned to the second caller comprises accessing, via a phone of the first caller, a contact list associated with the first caller, and determining, based on the contact list, that the second caller is associated with the unknown contact.
“ In typical use, the mobile device 114 uses the caller ID information in the local address book 102 to display information about calling and called parties to the user. When the mobile device 114 receives an inbound call, the mobile device determines the telephone number from which the inbound call was placed and accesses the local address book 102 to obtain caller ID information associated with the telephone number. For example, the mobile device 114 can display the name and picture of the caller to the user. If, however, the local address book 102 lacks information about the calling telephone number, the mobile device 114 typically displays only the telephone number to the user. ” (Devitt [Column 4 lines 49 to 60])
“When the mobile device 114 receives a call from an unknown caller (i.e., from a telephone number not listed in the local address book 102), the mobile application 104 queries the caller information server 110 for caller ID information associated with the telephone number that called the mobile device 114. The mobile application 104 receives the caller ID information from the caller information server 110 and presents the information to the mobile device 114 for display to the user. The mobile application 104 can also use the caller ID information to perform auxiliary services, such as blocking nuisance calls.” (Devitt [Column 5 lines 12 to 22])
It would have been obvious for someone of ordinary skill in the art before the effective filing
date of the claimed invention to have modified Efrati to incorporate the teachings of Devitt to include a contact list with the capabilities of providing a contact from the list and determine in a phone interface if the calling number is part of a contact list. The motivation and some of the benefits to include the contact list verification is discussed by Devitt in: “ Phone operators may provide caller ID information for an incoming or outgoing telephone call by matching a telephone number against a particular database. For example, a mobile or a landline telephone operator may maintain a database of user telephone numbers associated with a user-provided billing name. Thus, a telephone operator may provide the billing name associated with a particular telephone number. However, most databases associated with a mobile telephone operator are incomplete because telephone operators only have access to billing information associated with their own customers. In general, telephone operators do not share their customer billing information with their competitors. Additionally, the user-provided billing name may be an inaccurate description of a telephone number in a case where one billing name is provided for several associated telephone numbers such as employer-provided telephone plans or telephone numbers in a family or a group payment plan. Finally, users of prepaid phones generally do not provide a billing name to mobile phone operators, thus a vast number of billing names associated with mobile device phone numbers are not available to mobile telephone operators. In practice, major mobile telephone operators in the US generally do not provide caller ID services to its customers at this time.” (Devitt [Column 1 lines 35 to 57])
Claim(s) 9 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Efrati in view of Henrotte in further view of Shambaugh, Craig et al. (US 20030215066 A1) hereinafter Shambaugh.
Regarding claim 9, the rejection of claim 1 is incorporated, Efrati does not teach, but Shambaugh teaches:
(Currently Amended) The system of claim 1, wherein the caller is an automated caller and the input is text to speech machine generated input provided by a user or user group.
“One mode of practicing the invention is a method of automatic call handling allowing agent optimization in an automatic call distribution system that comprises synthesizing speech by using a script as input and generating speech as output, connecting a call from or to a call contact, and speaking to the call contact using speech generated using the prepared script as input.” (Shambaugh [0009]).
It would have been obvious for someone of ordinary skill in the art before the effective filing
date of the claimed invention to have modified Efrati to incorporate the teachings of Shambaugh to include instructions where the caller is an automated caller that follows the input that the user or group of users provide. The motivation and some of the benefits to include this modification to the previous invention is discussed by Shambaugh in the following sections: “Known systems for handling calls in an automated call center require an extensive staff of agents, because an individual agent must participate in each call. Such individual call handling can be less than optimally efficient if a number of different call contacts tend to pose the same questions and should receive the same responses from agents. In addition to inefficiency, the quality of individual call handling can be affected by individual agent fatigue or by inadvertent errors or omissions by individual agents.” (Shambaugh [0006]), and “Current technology only allows for one agent per call. This requires an agent to handle each call even if the same questions or responses are given to or from one call contact as another call contact. If the agent is consistently responding to many callers in the same way it would be desirable to automate the process and only require an agent to be involved when something unexpected or unique happens during the call.” (Shambaugh [0007]).
Regarding claim 19, the rejection of claim 11 is incorporated, furthermore arguments analogous to claim 9 are applicable.
Cited Reference Summary:
In this rejection the reference Efrati is relied upon for the framework of the contact or caller ID specific voice change, generating a synthetic speech which is a modified version of the original voice seed, storing the voice seed for the voice modification and retrieving the voice seed during the call. The reference Henrotte is relied upon for their voice masking system to protect the identity of the speaker. The reference Guedalia is relied upon because of their teachings contact list and contact selection. Reference Shambaugh is relied upon in this rejection for their teachings of interaction with an automated caller that reads a script based on the synthetic voice. The reference Devitt further expand in the use of a contact list and further helps identifying an unknown contact.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's
disclosure - see additional references cited on PTO-892.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HECTOR J. CRESPO FEBLES whose telephone number is (571)272-4512. The examiner can normally be reached Mon - Fri 7:30 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HECTOR J. CRESPO FEBLES/Examiner, Art Unit 2657
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657