Prosecution Insights
Last updated: October 02, 2026
Application No. 18/391,861

DEVICE, ENSEMBLE SYSTEM, SOUND REPRODUCING METHOD, AND NON-TRANSITORY COMPUTER-READABLE RECORDING MEDIUM

Non-Final OA §102§103
Filed
Dec 21, 2023
Priority
Dec 23, 2023 — nonprovisional of PCTJP2021023765
Examiner
GILLESPIE, NICOLE KATHLEEN
Art Unit
Tech Center
Assignee
Yamaha Corporation
OA Round
1 (Non-Final)
54%
Grant Probability
Moderate
1-2
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 54% of resolved cases
54%
Career Allowance Rate
36 granted / 66 resolved
-5.5% vs TC avg
Strong +50% interview lift
Without
With
+50.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
28 currently pending
Career history
74
Total Applications
across all art units

Statute-Specific Performance

§101
7.4%
-32.6% vs TC avg
§103
68.9%
+28.9% vs TC avg
§102
17.1%
-22.9% vs TC avg
§112
5.0%
-35.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 66 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 7 is objected to because of the following informalities: Regarding Claim 7, on line 1, change “claim 7” to –claim 6--. Appropriate correction is required. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1,3,4 and 6 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US20240112691 (Kumar), hereinafter US’691. Regarding Claim 1, US’691 discloses ‘A first device for a remote ensemble performed in a first venue and a second venue and provided in the first venue (US’691, ¶[0021]: multiple connected client devices 115a-115n communicating through network 105; ¶[0028]: Client device 115a generates a first audio stream that is transmitted through server 101 to another client device; ¶[0046]: The performance may be “music performed with instruments”; Fig. 5, uses first client device 510 and second client device 520; ¶[0099]:” a first audio stream of violin at the first client device 510 and a second audio stream of cello” being heard by the users as though performing in the same physical room), the first device comprising: a memory configured to store a performance sound estimation model (US’691, ¶[0033]: computing device 200 may be client device 115 and includes processor 235 and memory 237 ; ¶[0036]: memory 237 stores executable instructions; ¶[0051]:”the synthesizing machine-learning module 204 is stored in the memory 237 of the computing device 200; and ¶¶[0052]-[0054]:synthesizing ML model stored/ implemented by client computing device; ¶[0066]:that module trains/implements the ML model that synthesizes the future audio stream) ; and an estimation circuit configured to input, into the performance sound estimation model, a performance sound obtained by a second device provided in the second venue (US’691, Fig. 7 ¶¶[0102]-[0103]:“a first audio stream of a performance is received that is associated with a first client device 115” at the second client-device implementation; ¶¶[0048]-[0049]:The received stream is actual performance audio, the performance may be a person singing or a “a person playing a musical”; ¶¶[0061]-[0062]: the ML model then receives the audio stream itself “the bottleneck trunk 305 receives an audio stream generated by a client device 115 as input”, including the audio stream “as an input waveform”) to estimate an estimated future performance sound of the performance sound (US’691, ¶[0004]: synthesized audio “predicts a future of the performance based on audio features of the first audio stream”; ¶[0101]: DNN “extrapolates the audio stream into the future to complete the synthesis of the audio stream”), the performance sound estimation model comprising a trained model trained to learn a sound signal corresponding to the performance sound (US’691, ¶¶[0054]-0055]:trains ML model using audio-stream data supervised training uses “a training dataset with audio streams” the model’s layers learn increasingly detailed “features and patterns within the audio stream”; ¶[0071]: the VQ-VAE is trained using “using a training dataset that includes waveforms of audio streams “ and compares output waveforms during training) to estimate the estimated future performance sound based on the performance sound (US’691, ¶[0066]: model “model 300 learns the temporal mapping between the position in the audio stream and synthesizes future frames of the audio stream”, predicts the next frame based upon characteristics of the received stream and outputs synthesized audio incorporating those characteristics). Regarding Claim 3, US’691 discloses ‘An ensemble system for a remote ensemble performed in a first venue and a second venue (US’691, Fig. 5, “a first client device 510, a server 515 and a second client device 520”), the ensemble system comprising: a first terminal device provided in the first venue; and a second terminal device provided in the second venue (US’691, ¶[0099]:”a first audio stream of violin at the first client device 510 and a second audio stream of cello”, are heard by the associated users performing simultaneously as though they were “in the same physical room”), the first terminal device comprising: a first memory configured to store a second performance sound estimation model (US’691, Fig. 2, ¶¶[0075]-[0076]:computing device may be client device 115 and includes processor 235, memory 237, microphone 241, speaker 243 and storage 247, Fig. 2, ¶[0086]:Memory 237 stores instructions, and synthesizing ML module 204 is stored in memory 237); a first obtaining circuit configured to obtain a first performance sound generated in the first venue (US’691, ¶[0039]:”The microphone 241 includes hardware for detecting audio performed by a user 125,” including a user “playing the violin”; Fig. 5, ¶[0097]:”The first client device 510 receives an audio stream from the microphone 505); a first transmission circuit configured to transmit the first performance sound to the second terminal device (US’691, Fig. 5 depicts network data transmission between first client 510, server 515 and second client 520; ¶¶[0027]-[0029]:client devices communicate over network 105; client 115a generates a first audio stream, sends it to server 101, and the server transmits the communication to the other client); a first reception circuit configured to receive, from the second terminal device, a second performance sound generated in the second venue (US’691, ¶[0098]:”server 515 may receive a first stream from the first client device 510 and a second stream from the second client device 520”); a first estimation circuit configured to input, into the second performance sound estimation model, the second performance sound received at the first reception circuit to estimate an estimated future second performance sound of the second performance sound (US’691, Fig. 6, ¶¶[0100]-[0101]:the DNN “extrapolates the audio stream into the future to complete the synthesis of the audio stream”; ¶¶[0063]-[0066]:the ML model receives client-generated audio “as an input waveform” and ”the machine-learning model 300 learns the temporal mapping between the position in the audio stream and synthesizes future frames of the audio stream”); and a first sound outputting circuit configured to output the estimated future second performance sound (US’691, ¶[0040], ¶[0085]:the computing device includes speaker 243, which generates audio for playback from the combined/synthesized stream, Fig. 5 sends the combined stream to speaker 525 for playback), the second terminal device comprising: a second memory configured to store a first performance sound estimation model (US’691, ¶[0051] , ¶[0075]: computing device can be any client device 115, and its memory 237 stores the synthesizing ML module); a second obtaining circuit configured to obtain the second performance sound (US’691, ¶[0039]:”microphone 241 may detect a user 125 singing, user 125 playing the violin”; ¶[0046]:”the performance recognition module 202 receives the audio stream from the microphone 241”; ¶[0099]:”a first audio stream of violin at the first client device 510 and a second audio stream of cello are heard by associated users”); a second transmission circuit configured to transmit the second performance sound to the first terminal device (US’691, ¶[0098]:”the server 515 may receive a first stream from the first client device 510 and a second stream from the second client device 520”); a second reception circuit configured to receive the first performance sound from the first terminal device (US’691, Fig. 7 implements the synthesis method using a second client device 115 and “a first audio stream of a performance is received that is associated with a first client device 115”); a second estimation circuit configured to input, into the first performance sound estimation model, the first performance sound received at the second reception circuit to estimate an estimated future first performance sound of the first performance sound (US’691, Fig. 7, ¶[0102] and [0105]: places metaverse application 104 on the second client device and receives the first client’s performance stream; [0101]:the synthesis process maps the delayed audio and “is received by a deep neural network 635 (or other suitable model) that extrapolates the audio stream into the future”); and a second sound outputting circuit configured to output the estimated future first performance sound (US’691, ¶[0040]:”the speaker 243 includes hardware for generating audio for playback”), the first performance sound estimation model comprising a trained model trained to learn a first sound signal corresponding to the first performance sound to estimate the estimated future first performance sound based on the first performance sound (US’691, ¶[0058]:”The machine-learning model 300 includes …an audio waveform generator 315”; ¶[0066]:”the machine-learning model 300 learns the temporal mapping between the position in the audio stream and synthesizes future frames”), and the second performance sound estimation model comprising a trained model trained to learn a second sound signal corresponding to the second performance sound to estimate the estimated future second performance sound based on the second performance sound (US’691, ¶[0066]:”The machine-learning model 300 may predict the timing of the next frame in the audio stream based on the time offset, the rate of the audio stream, and a future time offset of the song based on the rate of the audio stream as compared to that of the reference audio and, as a result, output synthesized audio that encapsulates those features”). Regarding Claim 4, US’691 discloses ‘A sound reproducing method (US’691, ¶[0004]:”method to synthesize audio”) performed by a computer (¶[0033]: computing device may be client device 115 and includes processor 235, memory 237, microphone 241 and speaker 243, ¶[0039]: microphone detects “a user 125 playing the violin”) that is for a remote ensemble performed in a first venue and a second venue and that is provided in the first venue (US’691, ¶[0001]:”… physically in different locations but connected by a computer network…”; ¶[0079]:” a client device 115b performs both the synthesizing of the audio streams and the synchronizing of the audio streams “), the sound reproducing method comprising: inputting, into a performance sound estimation model, a performance sound obtained by a device provided in the second venue (US’691, ¶[0102]:Fig. 7 performs the future-audio synthesis at a second client device “the metaverse application 104 is stored on the second client device 115” … “a first audio stream of a performance is received that is associated with a first client device 115”; ¶[0079]: the synthesizing ML module receives time-stamped packets from the other client device and “a first audio stream of a performance is received that is associated with a first client device 115”) to estimate an estimated future performance sound of the performance sound, the performance sound estimation model (US’691, ¶[0101]: mapping supplied to DNN 635 which “extrapolates the audio stream into the future”; [Abstract]:the synthesized first audio stream “predicts a future of the performance based on audio features of the first audio stream”) comprising a trained model trained to learn a sound signal corresponding to the performance sound (US’691, ¶[0071]: trains VQ-VAE 355 and its codebook using “a training dataset that includes waveforms of audio streams“ and compares “output waveforms against input waveforms”) to estimate the estimated future performance sound based on the performance sound (US’691, ¶[0074]: the model receives the code vector derived from the input waveform and use transform layers to “extrapolate the code vectors over time”). Regarding Claim 6, US’691 discloses ‘A non-transitory computer-readable recording medium storing a program (US’691, [0116]:”a computer program may be stored in a non-transitory computer-readable storage medium”) that, when executed by at least one computer that is for a remote ensemble performed in a first venue and a second venue and that is provided in the first venue (US’691, ¶[0001]:”… physically in different locations but connected by a computer network…”; ¶[0079]:” a client device 115b performs both the synthesizing of the audio streams and the synchronizing of the audio streams “), cause the at least one computer to perform a method comprising: inputting, into a performance sound estimation model, a performance sound obtained by a device provided in the second venue to estimate an estimated future performance sound of the performance sound, the performance sound estimation model comprising a trained model trained to learn a sound signal corresponding to the performance sound to estimate the estimated future performance sound based on the performance sound. (Claim 6 corresponds to claim 4) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2,5 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over US’691, in view of US10944999/US20220303593 (Nicol), hereinafter US’999/ US’593. Regarding Claim 2, US’691 discloses ‘The first device according to claim 1, as discussed above. US’691 further discloses ‘wherein the performance sound estimation model is configured to learn a sound source corresponding to the performance sound (US’691, ¶[0058]:machine-learning model 300 includes a submodel decoder 310 “trained for a specific task; namely, to determine temporal correlations in the audio stream as compared to a reference audio”). US’691 does not expressly disclose the reference audio is a “rehearsal sound source”. However, US’999 discloses ‘rehearsal sound source (US’999, col. 2, lines 10-20: the server automatically processes a live musical performance using “reference audio data captured during a rehearsal”; col. 10, lines 52-57:” The reference audio data can include audio content recorded by microphones in a rehearsal…”; col. 10, lines 61-63:“Estimator 522 can instruct each player of a sound source at the performance …”) It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to configure the performance sound estimation model of US’691 to learn, as the sound source corresponding to the performance sound, a rehearsal sound source as taught by US’999, because US’999 teaches capturing reference audio from the corresponding performers/instruments during rehearsal and using that reference audio in processing the subsequent live performance. The modification would provide the reference audio representative of the sound sources participating in the performance. Regarding Claim 5, US’691 discloses ‘The sound reproducing method according to claim 4, as discussed above. ‘wherein the performance sound estimation model is configured to learn a rehearsal sound source corresponding to the performance sound. (Claim 5 corresponds to claim 2) Regarding Claim 7, US’691 discloses ‘The non-transitory computer-readable recording medium (US’691, [0116]:”a computer program may be stored in a non-transitory computer-readable storage medium”) according to claim 7, wherein the performance sound estimation model is configured to learn a rehearsal sound source corresponding to the performance sound. (Claim 7 corresponds to claim 2) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US10643593 (Kolen) teaches distributed virtual orchestra which different instrument simulators 302A-302N are hosted on different computing systems. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICOLE K GILLESPIE whose telephone number is (571)482-4187. The examiner can normally be reached Monday-Friday 7:30-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei K Hammond can be reached at (571)270-3819. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NICOLE K GILLESPIE/Examiner, Art Unit 2837 /DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837
Read full office action

Prosecution Timeline

Dec 21, 2023
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 9530178
NON-VOLATILE STORAGE FOR GRAPHICS HARDWARE
1y 7m to grant Granted Dec 27, 2016
Patent 9436740
VISUALIZATION OF CHANGING CONFIDENCE INTERVALS
4y 5m to grant Granted Sep 06, 2016
Patent 9437014
Method for Labeling Segments of Paths as Interior or Exterior
3y 1m to grant Granted Sep 06, 2016
Patent 9430851
Method for Converting Paths Defined by a Nonzero Winding Rule
3y 1m to grant Granted Aug 30, 2016
Patent 9400767
SUBGRAPH-BASED DISTRIBUTED GRAPH PROCESSING
2y 7m to grant Granted Jul 26, 2016
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
54%
Grant Probability
99%
With Interview (+50.3%)
3y 1m (~4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 66 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month