DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 5-6, 10-12, 14, 16-17, and 19 are rejected under 35 U.S.C. 102(a) (1) as being anticipated by Feinauer et al (“Feinauer”, US 20200193972).
Regarding Claim 1, Feinauer teaches a method, comprising: obtaining an audio stream from a participant device connected to a conference (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45),
wherein the audio stream represents speech of a user of the participant device (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45),
and wherein conference participants of the conference include the user and other participants (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45);
determining that an accent of the speech represented in the audio stream is different from accents of the other participants (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.);
receiving, from a device of one of the conference participants, a user request to modify the accent of the speech (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45);
modifying the accent of the speech in the audio stream to produce a modified audio stream (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45);
and causing output of the modified audio stream at the participant device from which the user request is received (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45).
Regarding Claim 5, Feinauer teaches the method of claim 1.
Feinauer further teaches wherein determining that the accent of the speech is different from accents of the other participants comprises: evaluating the accent of the speech using accent models (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.);
and identifying a difference between the evaluated accent and the accents associated with a plurality of other conference participants (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.).
Regarding Claim 6, Feinauer teaches the method of claim 1.
Feinauer further teaches wherein modifying the accent of the speech comprises: selecting an accent model based on a commonly detected accent among the other participants (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.);
and applying a filter to modify the speech such that the speech is modified to sound as if the speech were spoken in the selected accent model (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.).
Regarding Claim 10, Claim 10 is rejected with the same reasoning as Claim 1.
Regarding Claim 11, Feinauer teaches the system of claim 10.
Feinauer further teaches wherein the user request to modify the accent of the speech is received from the user (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.).
Regarding Claim 12, Feinauer teaches the system of claim 10.
Feinauer further teaches wherein the user request to modify the accent of the speech is received from one of the other participants (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants. Participant 240’s voice can also be modified.).
Regarding Claim 14, Feinauer teaches the system of claim 10.
wherein, to cause output of the modified audio stream, the one or more processors further configured to execute instructions stored in the one or more memories to: output the modified audio stream to a first subset of participant devices connected to the conference (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.),
wherein the first subset includes the participant device from which the user request is received (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.);
and output the audio stream to a second subset of participant devices connected to the conference (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants.).
Regarding Claim 16, Claim 16 is rejected with the same reasoning as Claim 1.
Regarding Claim 17, Feinauer teaches the one or more non-transitory computer readable media of claim 16.
Feinauer further teaches wherein the modified audio stream is produced and output during playback of a recording of the conference, and wherein the user request is initiated during the playback (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants. The user request is the user input of a voice in a different accent. The voice is recorded to be altered.).
Regarding Claim 19, Feinauer teaches the one or more non-transitory computer readable media of claim 16.
Feinauer further teaches wherein modifying the accent of the speech comprises: selecting a filter from an audio filter data store based on a modeled accent of the speech and a desired accent (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants. The filter is the chosen dialect such as an Australian accent.);
and applying the selected filter to the audio stream to modulate the accent to sound as if spoken in the desired accent (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The accent of call participant 200 is modified into an accent/dialect familiar to call participant 240 and therefore are different. Participant 240 can be multiple participants. The filter is the chosen dialect such as an Australian accent.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The test for obviousness is not whether the features of a secondary reference may be bodily incorporated into the structure of the primary reference; nor is it that the claimed invention must be expressly suggested in any one or all of the references. Rather, the test is what the combined teachings of the references would have suggested to those of ordinary skill in the art. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981).
Claims 3 is rejected under 35 U.S.C. 103 as being unpatentable over Feinauer in view of Li et al (“Li”, US 20110044324).
Regarding Claim 3, Feinauer teaches the method of claim 1.
Feinauer further teaches wherein modifying the accent of the speech in the audio stream comprises applying a filter specific to a speaker's selected accent (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45).
Feinauer does not explicitly teach applying a filter specific to a speaker's gender.
Li teaches applying a filter specific to a speaker's gender (par 39).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Feinauer with the tone modification of Li because it provides users to mask their voices, thereby further improving privacy.
Claims 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Feinauer in view of Chang et al (“Chang”, US 20140187210).
Regarding Claim 4, Feinauer teaches the method of claim 1.
Feinauer does not explicitly teach further comprising: identifying a second speech characteristic of the user; and maintaining the second speech characteristic unmodified during modification of the accent.
Chang teaches further comprising: identifying a second speech characteristic of the user (par 41; The accent or tempo of the speaker can be changed. Thus, the tempo is unmodified when an accent is modified.);
and maintaining the second speech characteristic unmodified during modification of the accent (par 41; The accent or tempo of the speaker can be changed. Thus, the tempo is unmodified when an accent is modified.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Feinauer with the voice modification of Chang because it allows for a speaker’s voice to modified for aesthetic or entertainment purposes (Chang; par 54).
Regarding Claim 13, Feinauer teaches the system of claim 10.
Feinauer does not explicitly teach wherein a cadence of the speech remains unmodified within the modified audio stream.
Chang teaches wherein a cadence of the speech remains unmodified within the modified audio stream (par 41; The accent or tempo of the speaker can be changed. Thus, the tempo (cadence) is unmodified when an accent is modified.);
and maintaining the second speech characteristic unmodified during modification of the accent (par 41; The accent or tempo of the speaker can be changed. Thus, the tempo (cadence) is unmodified when an accent is modified.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Feinauer with the voice modification of Chang because it allows for a speaker’s voice to modified for aesthetic or entertainment purposes (Chang; par 54).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Feinauer in view of Palamadai et al (“Palamadai”, US 20240097925).
Regarding Claim 20, Feinauer teaches the one or more non-transitory computer readable media of claim 16.
Feinauer further teaches participant device from which the user request is received (Fig. 2A, elements {200, 210, 220, 230, 240}, par 44; Fig. 2B, elements {200, 210, 220, 230, 240}, par 45; The user request is the user input of a voice in a different accent. The voice is recorded to be altered).
Feinauer does not explicitly teach the operations further comprising: storing configuration data indicative of a filter used to modify the accent at the participant device; and automatically applying the filter during a future conference when the user of the participant device providing the audio stream is a participant.
Palamadai teaches the operations further comprising: storing configuration data indicative of a filter used to modify the accent at the participant device (par 90);
and automatically applying the filter during a future conference when the user of the participant device providing the audio stream is a participant (par 90).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Feinauer with the filtering component of Palamadai because it allows for conferencing participants to change how they look, as well additional ways to change how they sound (Palamadai; par 90), thereby improving collaboration and engagement.
Allowable Subject Matter
Claims 2, 7-9, 15, and 18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
In interpreting the currently amended claims, in light of the specification, the Examiner finds the claimed invention to be patentably distinct from the prior art of record.
Regarding Claim 2, the closest prior art of record Feinauer et al (US 20200193972) in view of Chang et al (US 20140187210) in further view of Palamadai et al (US 20240097925) and in even further view of Li et al (US 20110044324) does not teach a method, comprising: obtaining an audio stream from a participant device connected to a conference, wherein the audio stream represents speech of a user of the participant device, and wherein conference participants of the conference include the user and other participants; determining that an accent of the speech represented in the audio stream is different from accents of the other participants; receiving, from a device of one of the conference participants, a user request to modify the accent of the speech; modifying the accent of the speech in the audio stream to produce a modified audio stream; and causing output of the modified audio stream at the participant device from which the user request is received; further comprising: in response to determining that the accent is different, presenting a prompt at the device of one of the conference participants to recommend modifying the accent.
Regarding Claim 7, the closest prior art of record Feinauer et al (US 20200193972) in view of Chang et al (US 20140187210) in further view of Palamadai et al (US 20240097925) and in even further view of Li et al (US 20110044324) does not teach a method, comprising: obtaining an audio stream from a participant device connected to a conference, wherein the audio stream represents speech of a user of the participant device, and wherein conference participants of the conference include the user and other participants; determining that an accent of the speech represented in the audio stream is different from accents of the other participants; receiving, from a device of one of the conference participants, a user request to modify the accent of the speech; modifying the accent of the speech in the audio stream to produce a modified audio stream; and causing output of the modified audio stream at the participant device from which the user request is received; further comprising: identifying a pitch of the speech from the audio stream; and causing a prompt to be displayed to a participant when the pitch is determined to be outside a defined threshold range.
Regarding Claim 8, the closest prior art of record Feinauer et al (US 20200193972) in view of Chang et al (US 20140187210) in further view of Palamadai et al (US 20240097925) and in even further view of Li et al (US 20110044324) does not teach a method, comprising: obtaining an audio stream from a participant device connected to a conference, wherein the audio stream represents speech of a user of the participant device, and wherein conference participants of the conference include the user and other participants; determining that an accent of the speech represented in the audio stream is different from accents of the other participants; receiving, from a device of one of the conference participants, a user request to modify the accent of the speech; modifying the accent of the speech in the audio stream to produce a modified audio stream; and causing output of the modified audio stream at the participant device from which the user request is received; wherein modifying the accent of the speech comprises: altering an audio file corresponding to the audio stream stored in a recording of the conference; and pausing playback of the recording while the audio file is altered.
Regarding Claim 9, the closest prior art of record Feinauer et al (US 20200193972) in view of Chang et al (US 20140187210) in further view of Palamadai et al (US 20240097925) and in even further view of Li et al (US 20110044324) does not teach a method, comprising: obtaining an audio stream from a participant device connected to a conference, wherein the audio stream represents speech of a user of the participant device, and wherein conference participants of the conference include the user and other participants; determining that an accent of the speech represented in the audio stream is different from accents of the other participants; receiving, from a device of one of the conference participants, a user request to modify the accent of the speech; modifying the accent of the speech in the audio stream to produce a modified audio stream; and causing output of the modified audio stream at the participant device from which the user request is received; wherein receiving the user request to modify the accent of the speech comprises: presenting a prompt to the user, wherein the prompts recommends accent modification based on detected accent difference.
Regarding Claim 15, the closest prior art of record Feinauer et al (US 20200193972) in view of Chang et al (US 20140187210) in further view of Palamadai et al (US 20240097925) and in even further view of Li et al (US 20110044324) does not teach a system, comprising: one or more memories; and one or more processors configured to execute instructions stored in the one or more memories to: obtain an audio stream from a participant device connected to a conference, wherein the audio stream represents speech of a user of the participant device, and wherein conference participants of the conference include the user and other participants; determine that an accent of the speech represented in the audio stream is different from accents of the other participants; receive, from a device of one of the conference participants, a user request to modify the accent of the speech; modify the accent of the speech in the audio stream to produce a modified audio stream; and cause output of the modified audio stream at the participant device from which the user request is received; wherein the one or more processors further configured to execute instructions stored in the one or more memories to: transmit a notification to the participant device providing the audio stream indicating that the accent of the speech is being modified for one or more other participants; and specifying, within the notification, an alteration of the accent.
Regarding Claim 18, the closest prior art of record Feinauer et al (US 20200193972) in view of Chang et al (US 20140187210) in further view of Palamadai et al (US 20240097925) and in even further view of Li et al (US 20110044324) does not teach one or more non-transitory computer readable media storing instructions that, when executed by one or more processors, perform operations comprising: obtaining an audio stream from a participant device connected to a conference, wherein the audio stream represents speech of a user of the participant device, and wherein conference participants of the conference include the user and other participants; determining that an accent of the speech represented in the audio stream is different from accents of the other participants; receiving, from a device of one of the conference participants, a user request to modify the accent of the speech; modifying the accent of the speech in the audio stream to produce a modified audio stream; and causing output of the modified audio stream at the participant device from which the user request is received; wherein modifying the accent of the speech comprises: altering an audio file corresponding to the audio stream stored in a recording of the conference; and pausing playback of the recording while the audio file is altered.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Singh et al (US 20020126626), Abstract - A conferencing server is provided for a data network telephony system for facilitating multi-party call conferencing. The conferencing server generally includes a session initiation protocol (SIP) signaling interface and a media conferencing module. The media conferencing module includes a number of selectable media decoders compliant with known media CODEC protocols, such as G.711. A number of media stream queues s are selectively coupled to the media decoders for storing the decoded incoming media streams. A jitter correction processor is provided to compensate for arrival time jitter in the data stored in the media stream queues. A mixer, which receives the jitter corrected data from each of the queues, generates an aggregate conferencing stream of all active participants and generates individual participant conference streams for each active participant in the conference. A number of selectable media encoders are then used to encode the individual participant conference streams in accordance with a media CODEC protocol supported by the respective participant. The encoded participant conference streams are then distributed to the various conference participants via the data network, such as the Internet.
Yu (US 20230412656), Abstract - Aspect ratios used to display video streams within a graphical user interface (GUI) of a video conference are dynamically adjusted based on events detected during the video conference. According to one approach, during a video conference between a first device and a second device, a first video stream from the first device and a second video stream from the second device are both displayed within the GUI using an initial aspect ratio. Based on the first video stream, an event corresponding to a change in a number of people participating in the video conference from the first device is determined. Based on the event, an adjusted aspect ratio to use for displaying the first video stream within the GUI is determined. The first video stream is displayed within the GUI using the adjusted aspect ratio while the second video stream remains displayed within the GUI using the initial aspect ratio.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RAQIUL AMIN CHOUDHURY whose telephone number is (571)272-2482. The examiner can normally be reached Monday-Friday 7:30 AM - 5:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John Follansbee can be reached at 571-272-3964. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RAQIUL A CHOUDHURY/Examiner, Art Unit 2444